Qwen3.8 Flash Next leads the local field at 96.0 combined. Qwen3.8 27B is second at 95.0. Fastest decode: Qwen3.6 35B at 162 tok/s.
Cloud reference (not comparable): gpt-6-astra · medium scored 100.0 on the same cases via Codex CLI.
All scored models, click a column to sort. Cloud models ran through Codex CLI — different harness, tokenizer and reasoning accounting — so they are ranked here flagged ☁ as not comparable; their detail tables stay in the cloud band below.
| # | Model | Combined | Single | Agentic | Tokens | Run time | TG tok/s | Memory GiB |
|---|---|---|---|---|---|---|---|---|
| 1 | gpt-6-astra · medium CloudCloud leader |
100.0 |
100.0 |
100.0 |
27,503 | 29.5 min | — | — |
| 2 | gpt-6-sol · medium CloudCloud · Codex CLI |
99.4 |
99.3 |
100.0 |
40,893 | 31.1 min | — | — |
| 3 | Flash Next Jundot/Qwen3.8-Flash-Next-oQ4e-mtp Combined leader |
96.0 |
95.3 |
100.0 |
248,552 | 1.0 h | 75 | 69.0 |
| 4 | Qwen3.8 27B Jundot/Qwen3.8-27B-oQ8e-mtp Complete |
95.0 |
94.6 |
97.7 |
263,606 | 3.3 h | 46 | 28.1 |
| 5 | gpt-5.6-terra · medium CloudCloud · Codex CLI |
94.3 |
93.5 |
99.1 |
36,098 | 25.3 min | — | — |
| 6 | gpt-oss 120B mlx-community/gpt-oss-120b-MXFP4-Q8 Complete |
92.2 |
92.4 |
91.1 |
183,174 | 1.3 h | 90 | 59.0 |
| 7 | Gemma 4 26B Jundot/gemma-4-26B-A4B-it-oQ6 Complete |
91.8 |
92.0 |
90.2 |
510,477 | 3.7 h | 95 | 20.7 |
| 8 | Gemma 4 31B Jundot/gemma-4-31B-it-oQ6e-mtp Complete |
89.4 |
89.3 |
90.0 |
54,108 | 1.6 h | 19 | 26.0 |
| 9 | Gemma 4 12B mlx-community/gemma-4-12B-it-8bit Complete |
87.8 |
88.8 |
81.5 |
553,162 | 8.4 h | 39 | 12.2 |
| 10 | gpt-oss 20B mlx-community/gpt-oss-20b-MXFP4-Q8 Complete |
85.9 |
89.1 |
66.3treat as unpatched |
240,338 | 1.2 h | 136 | 10.9 |
| 11 | gpt-6-luna · medium CloudCloud · Codex CLI |
84.5 |
81.9 |
100.0 |
22,956 | 18.1 min | — | — |
| 12 | Qwen3.6 35B Jundot/Qwen3.6-35B-A3B-oQ6-fp16-mtp Complete |
70.8 |
67.4 |
91.2 |
895,719 | 3.8 h | 162 | 28.3 |
| 13 | Laguna S 2.1 mlx-community/Laguna-S-2.1-oQ5e Complete |
67.4 |
66.7 |
71.1 |
233,792 | 5.7 h | 48 | 72.9 |
| 14 | GLM-4.7 Flash lmstudio-community/GLM-4.7-Flash-MLX-8bit Lowest combined |
55.8 |
52.1 |
77.7 |
542,137 | 9.1 h | 71 | 29.6 |
Every model in this benchmark, as configured on this machine.
Mean score per category, local models. The leader in each glows.
Mean of two seeds per case. Click any cell to open that case in the explorer.
| Model | N1 | N2 | N3 | N4 | N5 | C2 | C3 | C4 | J1 | J2 | S1 | S2 | R1 | R2 | L1 | L5 | L3 | M1 | M2 | M3 | M4 | M5 | PL1 | PL2 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Flash Next | ||||||||||||||||||||||||
| Qwen3.8 27B | ||||||||||||||||||||||||
| gpt-oss 120B | ||||||||||||||||||||||||
| Gemma 4 26B | ||||||||||||||||||||||||
| Gemma 4 31B | ||||||||||||||||||||||||
| Gemma 4 12B | ||||||||||||||||||||||||
| gpt-oss 20B | ||||||||||||||||||||||||
| Qwen3.6 35B | ||||||||||||||||||||||||
| Laguna S 2.1 | ||||||||||||||||||||||||
| GLM-4.7 Flash |
The saved answer for every model, per seed — prompts, scores, token counts, grader detail.
Four multi-turn tool-using tasks, two seeds each, graded by hidden unit tests in the task repos.
Hidden tests; an untouched repo scores 5/14.
| Model | Score | Tests | Max run | Notes |
|---|---|---|---|---|
| Flash Next | 100% | 28/28 | 1 min | all hidden tests, both seeds |
| Qwen3.8 27B | 100% | 28/28 | 2 min | all hidden tests, both seeds |
| gpt-oss 120B | 96% | 27/28 | 1 min | |
| Gemma 4 26B | 100% | 28/28 | 1 min | all hidden tests, both seeds |
| Gemma 4 31B | 93% | 26/28 | 3 min | |
| Gemma 4 12B | 68% | 19/28 | 31 min | hit 30-min limit |
| gpt-oss 20B | 89% | 25/28 | 3 min | |
| Qwen3.6 35B | 93% | 26/28 | 1 min | |
| Laguna S 2.1 | 100% | 28/28 | 4 min | all hidden tests, both seeds |
| GLM-4.7 Flash | 86% | 24/28 | 2 min |
Hidden tests; an untouched repo scores 1/14.
| Model | Score | Tests | Max run | Notes |
|---|---|---|---|---|
| Flash Next | 100% | 28/28 | 3 min | all hidden tests, both seeds |
| Qwen3.8 27B | 100% | 28/28 | 5 min | all hidden tests, both seeds |
| gpt-oss 120B | 93% | 26/28 | 12 min | hit 40-turn limit |
| Gemma 4 26B | 93% | 26/28 | 6 min | |
| Gemma 4 31B | 93% | 26/28 | 3 min | |
| Gemma 4 12B | 93% | 26/28 | 20 min | |
| gpt-oss 20B | 50% | 14/28 | 3 min | repo unchanged |
| Qwen3.6 35B | 79% | 22/28 | 9 min | |
| Laguna S 2.1 | 50% | 14/28 | 34 min | repo unchangedhit 30-min limit |
| GLM-4.7 Flash | 79% | 22/28 | 35 min | hit 30-min limithit 40-turn limit |
Hidden tests; an untouched repo scores 1/14.
| Model | Score | Tests | Max run | Notes |
|---|---|---|---|---|
| Flash Next | 100% | 28/28 | 2 min | all hidden tests, both seeds |
| Qwen3.8 27B | 100% | 28/28 | 9 min | all hidden tests, both seeds |
| gpt-oss 120B | 100% | 28/28 | 2 min | all hidden tests, both seeds |
| Gemma 4 26B | 93% | 26/28 | 6 min | |
| Gemma 4 31B | 93% | 26/28 | 7 min | |
| Gemma 4 12B | 96% | 27/28 | 32 min | hit 30-min limit |
| gpt-oss 20B | 82% | 23/28 | 5 min | hit 40-turn limit |
| Qwen3.6 35B | 96% | 27/28 | 3 min | |
| Laguna S 2.1 | 100% | 28/28 | 26 min | all hidden tests, both seeds |
| GLM-4.7 Flash | 71% | 20/28 | 21 min | hit 40-turn limit |
Hidden tests; an untouched repo scores 0/1.
| Model | Score | Tests | Max run | Notes |
|---|---|---|---|---|
| Flash Next | 100% | 32/32 | 6 min | all hidden tests, both seeds |
| Qwen3.8 27B | 91% | 29/32 | 11 min | |
| gpt-oss 120B | 75% | 24/32 | 3 min | load-sensitive test failed |
| Gemma 4 26B | 75% | 24/32 | 34 min | hit 30-min limitload-sensitive test failed |
| Gemma 4 31B | 81% | 26/32 | 8 min | |
| Gemma 4 12B | 69% | 22/32 | 44 min | hit 30-min limitload-sensitive test failed |
| gpt-oss 20B | 44% | 14/17 | 6 min | hit 40-turn limit |
| Qwen3.6 35B | 97% | 31/32 | 20 min | load-sensitive test failed |
| Laguna S 2.1 | 34% | 11/17 | 34 min | repo unchangedhit 30-min limitload-sensitive test failed |
| GLM-4.7 Flash | 75% | 24/32 | 25 min | hit 40-turn limitload-sensitive test failed |
Completion tokens across the single-turn answers, and what thinking cost.
All ran at max_tokens 32,768; OpenCode's own per-model output limits listed last. "Score w/ thinking" is "—" for models whose records contain no reasoning.
| Model | Total tokens | Median | Max | Hit 32K | >16K | Visible thinking | Score w/ thinking | Score direct | OpenCode limit |
|---|---|---|---|---|---|---|---|---|---|
| Qwen3.6 35B | 895,719 | 14,704.5 | 32,768 | 13 | 22 | 48 | 67.4 | — | 32,000 |
| Gemma 4 12B | 553,162 | 10,503.0 | 32,768 | 2 | 9 | 48 | 88.8 | — | 16,384 |
| GLM-4.7 Flash | 542,137 | 8,512.5 | 32,768 | 7 | 10 | 48 | 52.1 | — | 16,384 |
| Gemma 4 26B | 510,477 | 9,885.0 | 32,768 | 1 | 5 | 48 | 92.0 | — | 16,384 |
| Qwen3.8 27B | 263,606 | 3,733.0 | 19,833 | 0 | 2 | 48 | 94.6 | — | 32,000 |
| Flash Next | 248,552 | 3,924.5 | 24,812 | 0 | 3 | 48 | 95.3 | — | 32,000 |
| gpt-oss 20B | 240,338 | 3,516.0 | 32,611 | 0 | 2 | 48 | 89.1 | — | 16,384 |
| Laguna S 2.1 | 233,792 | 703.0 | 32,768 | 2 | 5 | 16 | 68.0 | 66.1 | 32,000 |
| gpt-oss 120B | 183,174 | 2,583.0 | 20,543 | 0 | 2 | 48 | 92.4 | — | 16,384 |
| Gemma 4 31B | 54,108 | 803.0 | 8,969 | 0 | 0 | 0 | — | 89.3 | 16,384 |
Clean single-workload measurements (cache-busted prefill, 512-token decode). Prefill columns are prompts sized by characters, not tokens — the x value differs per tokenizer.
| Model | PP @2K chars tok/s | PP @8K chars | PP @32K chars | Decode tok/s |
|---|---|---|---|---|
| Qwen3.6 35B | 2,915 | 4,669 | 4,496 | 162 |
| gpt-oss 20B | 2,561 | 4,201 | 4,304 | 136 |
| Gemma 4 26B | 2,076 | 3,232 | 3,047 | 95 |
| gpt-oss 120B | 1,198 | 1,909 | 1,981 | 90 |
| Flash Next | 1,218 | 2,177 | 2,460 | 75 |
| GLM-4.7 Flash | 1,725 | 2,710 | 1,702 | 71 |
| Laguna S 2.1 | 813 | 1,109 | 1,030 | 48 |
| Qwen3.8 27B | 567 | 714 | 710 | 46 |
| Gemma 4 12B | 1,037 | 1,654 | 1,581 | 39 |
| Gemma 4 31B | 382 | 567 | 558 | 19 |
One point per finished local model: combined score against clean decode speed. The top-right corner is the sweet spot — high quality at high tokens per second. Cloud models never enter the charts.
x = each model's first clean speed.py decode run (512 tokens, cache-busted); y = combined over all cases and tasks. Every finished local model has a clean speed run.
Cache-busted prompts grown toward 10K–200K tokens, single-token prefill, repetitions averaged; the x-axis is measured prompt_tokens, so curves stay honest across tokenizers. Collected by the speed.py pp_ctx stage; models appear once swept. The sweep postdates the table's first-run figures, so values may come from a different session.
Cache-busted prompts grown toward 10K–200K tokens, 512-token decode, repetitions averaged; the x-axis is measured prompt_tokens, so curves stay honest across tokenizers. Collected by the speed.py tg_ctx stage; models appear once swept. The sweep postdates the table's first-run figures, so values may come from a different session.
The OpenCode resident model stays loaded; loading a benchmark model evicts others.
After a run, reload the resident model (after-run.sh).
Run through codex exec on the user's ChatGPT Plus plan. Different harness, tokenizer and
reasoning accounting than the local oMLX runs — scores are indicative, not directly comparable.
| Model | Combined | Single | Agentic | Strict /48 | Tool answers | Peeked | 30-min timeouts | Plan usage |
|---|---|---|---|---|---|---|---|---|
| gpt-6-astra · medium | 100.0 | 100.0 | 100.0 | 48/48 | 0 | 0 | 0 | 5-hour 99% weekly 45% |
| gpt-6-sol · medium | 99.4 | 99.3 | 100.0 | 44/48 | 0 | 0 | 0 | 5-hour 16% weekly 32% |
| gpt-5.6-terra · medium | 94.3 | 93.5 | 99.1 | 43/48 | 0 | 0 | 0 | 5-hour 32% weekly 35% |
| gpt-6-luna · medium | 84.5 | 81.9 | 100.0 | 30/48 | 0 | 0 | 0 | 5-hour 1% weekly 30% |
| Model | N1 | N2 | N3 | N4 | N5 | C2 | C3 | C4 | J1 | J2 | S1 | S2 | R1 | R2 | L1 | L5 | L3 | M1 | M2 | M3 | M4 | M5 | PL1 | PL2 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| gpt-6-astra · medium | ||||||||||||||||||||||||
| gpt-6-sol · medium | ||||||||||||||||||||||||
| gpt-5.6-terra · medium | ||||||||||||||||||||||||
| gpt-6-luna · medium |
24 single-turn cases (Python, Java, SQL, trace debugging, reasoning, code reading) × 2 seeds (500, 501), plus 4 multi-turn tool-using agent tasks (A1–A4) × 2 seeds graded by hidden unit tests in the task repos. Partial credit throughout; combined = mean over all 24 cases and 4 tasks, each weighted equally after averaging seeds, so agent work is 4/28 of it. An untouched repo scores A1 5/14, A2 1/14, A3 1/14, A4 0/1.
Thinking was on for every model except gemma-4-31B-it-oQ6e-mtp: its 48 single-turn rows and 8 agent rows went out with empty sampling and no reasoning (the X-1 finding), so its score is a thinking-off run and is disclosed as such. Per-model expectations now live in models.json: `request_extra` sends `enable_thinking` per request for the Gemma 4 models (whose chat template defaults thinking off), and `thinking_default` marks models that reason by server default. run.py merges `request_extra` into every request and fails the warmup when a model expected to think returns no reasoning. Sampling came from ~/.omlx/model_settings.json per model. Derived from the records: Gemma 4 31B returned no reasoning on any row — it ran with thinking off.
oMLX 0.7.0rc1 has a Harmony parser bug that silently drops completions whose first message is a tool call (empty stop). The harness resamples blank replies up to 16 times, but on the unpatched build gpt-oss-20b still lost whole agent runs to it. Agentic scores reflect oMLX + model, not the model alone.gpt-oss 20B’s agent runs are recorded as patched with ../patches/omlx-0.7.0rc1-harmony-tool-calls.patch; patch state unconfirmed — the installed bundle checks UNPATCHED; treat all agent runs as unpatched. The runs logged as unpatched for gpt-oss 20B needed 91 blank-reply resamples, 4 of 8 runs were blocked and agentic was 49.3%; the reruns recorded as patched logged 4, with 0 blocked, for 66.3%. Patch state is unconfirmed — treat every agent run as unpatched. The replaced runs stay in agent-results.jsonl and agent-transcripts-superseded/.
A4's test_amortized_constant_time and test_idle_keys_are_forgotten fail on a busy machine. a4_idle_regrade.py rebuilds saved A4 diffs and re-grades them on an idle machine (state/a4-idle-regrade.json); those rechecks are reported separately, and the tables keep the live grades. After the run, 16 A4 runs were rebuilt from their stored diffs and re-graded on the idle machine; every one reproduced its live score.
Agent runs stop at 30 minutes or 40 turns; Laguna hit the clock on A2 and A4 both seeds. Loading a benchmark model evicts others on this 128 GiB M5 Max — Flash Next alone stays resident at ~69.5 GiB, and OpenCode traffic during a run contaminates speed numbers, so speed was measured in a dedicated clean pass.
The same cases and tasks ran through Codex CLI on a ChatGPT Plus plan. Cloud results are reported only in their own band: different harness, different tokenizer, reasoning tokens accounted differently, and every run spends real plan quota.
v1 (thinking off, 2,000 tokens) recommended switching build/plan to Qwen3.6-35B. With thinking on and a 32K budget, 35B truncated 13 of 48 answers and fell to 70.8 — that recommendation is withdrawn. v1 and v2 scores are not comparable.