Local LLM benchmark · v2 report built 2026-10-06 06:45

Which large model should
run OpenCode?

Qwen3.8 Flash Next leads the local field at 96.0 combined. Qwen3.8 27B is second at 95.0. Fastest decode: Qwen3.6 35B at 162 tok/s.
Cloud reference (not comparable): gpt-6-astra · medium scored 100.0 on the same cases via Codex CLI.

10 local models24 cases × 2 seeds4 agent tasksoMLX 0.7.0Apple M5 Max · 128 GiBmax_tokens 32,768seeds 500–501
1Flash Next
0
combined
single
95.3
agentic
100.0
75 tok/s TG 69.0 GiB loaded 0 hit 32K43/48 strict
2Qwen3.8 27B
0
combined
single
94.6
agentic
97.7
46 tok/s TG 28.1 GiB loaded 0 hit 32K40/48 strict
3gpt-oss 120B
0
combined
single
92.4
agentic
91.1
90 tok/s TG 59.0 GiB loaded 0 hit 32K38/48 strict
☁gpt-6-astra · medium
100.0
combined · cloud
Codex CLI, different harness and tokenizer. Shown for scale only — see the cloud band below.

League table

All scored models, click a column to sort. Cloud models ran through Codex CLI — different harness, tokenizer and reasoning accounting — so they are ranked here flagged ☁ as not comparable; their detail tables stay in the cloud band below.

#ModelCombinedSingleAgenticTokensRun timeTG tok/sMemory GiB
1
gpt-6-astra · medium
Cloud
Cloud leader
100.0
100.0
100.0
27,50329.5 min — —
2
gpt-6-sol · medium
Cloud
Cloud · Codex CLI
99.4
99.3
100.0
40,89331.1 min — —
3
Flash Next
Jundot/Qwen3.8-Flash-Next-oQ4e-mtp
Combined leader
96.0
95.3
100.0
248,5521.0 h 75 69.0
4
Qwen3.8 27B
Jundot/Qwen3.8-27B-oQ8e-mtp
Complete
95.0
94.6
97.7
263,6063.3 h 46 28.1
5
gpt-5.6-terra · medium
Cloud
Cloud · Codex CLI
94.3
93.5
99.1
36,09825.3 min — —
6
gpt-oss 120B
mlx-community/gpt-oss-120b-MXFP4-Q8
Complete
92.2
92.4
91.1
183,1741.3 h 90 59.0
7
Gemma 4 26B
Jundot/gemma-4-26B-A4B-it-oQ6
Complete
91.8
92.0
90.2
510,4773.7 h 95 20.7
8
Gemma 4 31B
Jundot/gemma-4-31B-it-oQ6e-mtp
Complete
89.4
89.3
90.0
54,1081.6 h 19 26.0
9
Gemma 4 12B
mlx-community/gemma-4-12B-it-8bit
Complete
87.8
88.8
81.5
553,1628.4 h 39 12.2
10
gpt-oss 20B
mlx-community/gpt-oss-20b-MXFP4-Q8
Complete
85.9
89.1
66.3
treat as unpatched
240,3381.2 h 136 10.9
11
gpt-6-luna · medium
Cloud
Cloud · Codex CLI
84.5
81.9
100.0
22,95618.1 min — —
12
Qwen3.6 35B
Jundot/Qwen3.6-35B-A3B-oQ6-fp16-mtp
Complete
70.8
67.4
91.2
895,7193.8 h 162 28.3
13
Laguna S 2.1
mlx-community/Laguna-S-2.1-oQ5e
Complete
67.4
66.7
71.1
233,7925.7 h 48 72.9
14
GLM-4.7 Flash
lmstudio-community/GLM-4.7-Flash-MLX-8bit
Lowest combined
55.8
52.1
77.7
542,1379.1 h 71 29.6

Specs and settings

Every model in this benchmark, as configured on this machine.

Qwen3.6 35B
oQ6 · fp16 · MTP
70.8
Model fileJundot/Qwen3.6-35B-A3B-oQ6-fp16-mtp
Size35B
Active3B
ArchitectureMoE · 256 experts, 8 routed + 1 shared · 40 layers · hybrid attention
Samplingtemperature 0.6 · top_p 0.95 · top_k 20 · min_p 0.0 · repetition_penalty 1.0
Rolesassistant
Run time3.8 h · thinking on
gpt-oss 20B
MXFP4 · Q8
85.9
Model filemlx-community/gpt-oss-20b-MXFP4-Q8
Size20B
Active3.6B
ArchitectureMoE · Harmony chat format
Samplingtemperature 1.0 · top_p 1.0
Roles
Run time1.2 h · thinking on
gpt-oss 120B
MXFP4 · Q8
92.2
Model filemlx-community/gpt-oss-120b-MXFP4-Q8
Size120B
Active5.1B
ArchitectureMoE · Harmony chat format
Samplingtemperature 1.0 · top_p 1.0
Roles
Run time1.3 h · thinking on
Laguna S 2.1
oQ5e
67.4
Model filemlx-community/Laguna-S-2.1-oQ5e
Size70B
Active8B
ArchitectureMoE
Samplingtemperature 0.7 · top_p 0.95 · top_k 20 · min_p 0.0 · repetition_penalty 1.0
Roles
Run time5.7 h · thinking on
Qwen3.8 27B
oQ8e · MTP
95.0
Model fileJundot/Qwen3.8-27B-oQ8e-mtp
Size27B
Active27B
ArchitectureDense · 64 layers · hybrid attention
Samplingtemperature 0.6 · top_p 0.95 · top_k 20 · min_p 0.0 · repetition_penalty 1.0
Rolesdebugger
Run time3.3 h · thinking on
Flash Next
oQ4e · MTP
96.0
Model fileJundot/Qwen3.8-Flash-Next-oQ4e-mtp
Size125B + 51B n-gram
Active6B
ArchitectureMoE · 512 experts, 10 routed + 1 shared · 48 layers · n-gram tables on SSD
Samplingtemperature 0.6 · top_p 0.95 · top_k 20 · min_p 0.0 · repetition_penalty 1.0
Rolesdefault · build · plan · explore · review
Run time1.0 h · thinking on
GLM-4.7 Flash
8-bit
55.8
Model filelmstudio-community/GLM-4.7-Flash-MLX-8bit
Size30B
Active3B
ArchitectureMoE · 64 experts, 4 routed + 1 shared · 47 layers · MLA attention
Samplingtemperature 0.7 · top_p 1.0 · min_p 0.0 · repetition_penalty 1.0
Roles
Run time9.1 h · thinking on
Gemma 4 26B
oQ6
91.8
Model fileJundot/gemma-4-26B-A4B-it-oQ6
Size26B
Active3.8B
ArchitectureMoE · 128 experts, 8 routed · 30 layers · 5 of 6 sliding-window
Samplingtemperature 1.0 · top_p 0.95 · top_k 64
Roles
Run time3.7 h · thinking on
Gemma 4 12B
8-bit
87.8
Model filemlx-community/gemma-4-12B-it-8bit
Size12B
Active12B
ArchitectureDense · 48 layers · 5 of 6 sliding-window
Samplingtemperature 1.0 · top_p 0.95 · top_k 64
Roles
Run time8.4 h · thinking on
Gemma 4 31B
oQ6e
89.4
Model fileJundot/gemma-4-31B-it-oQ6e-mtp
Size31b
Active—
Architecture—
Samplingserver defaults
Roles
Run time1.6 h · thinking off

By category

Mean score per category, local models. The leader in each glows.

Python

Flash Next
99%
Qwen3.8 27B
97%
gpt-oss 20B
96%
Gemma 4 31B
95%
gpt-oss 120B
94%
Laguna S 2.1
92%
Gemma 4 12B
92%
Gemma 4 26B
92%
Qwen3.6 35B
82%
GLM-4.7 Flash
69%

Java

Qwen3.8 27B
100%
Flash Next
100%
Gemma 4 31B
99%
gpt-oss 120B
84%
Gemma 4 12B
72%
Gemma 4 26B
67%
Qwen3.6 35B
63%
gpt-oss 20B
62%
Laguna S 2.1
54%
GLM-4.7 Flash
32%

SQL

gpt-oss 20B
100%
gpt-oss 120B
100%
Qwen3.8 27B
100%
Flash Next
100%
Gemma 4 26B
100%
Gemma 4 12B
100%
Gemma 4 31B
100%
Qwen3.6 35B
75%
Laguna S 2.1
50%
GLM-4.7 Flash
0%

Trace

Flash Next
99%
gpt-oss 120B
98%
Qwen3.8 27B
95%
Gemma 4 26B
95%
gpt-oss 20B
87%
Gemma 4 12B
77%
Gemma 4 31B
70%
Laguna S 2.1
50%
Qwen3.6 35B
47%
GLM-4.7 Flash
14%

Reasoning

Gemma 4 26B
94%
gpt-oss 120B
88%
Qwen3.8 27B
88%
Flash Next
88%
Gemma 4 12B
88%
gpt-oss 20B
85%
Gemma 4 31B
81%
Qwen3.6 35B
62%
GLM-4.7 Flash
56%
Laguna S 2.1
48%

Code reading

gpt-oss 120B
100%
Qwen3.8 27B
100%
Flash Next
100%
Gemma 4 26B
98%
Gemma 4 12B
98%
Gemma 4 31B
97%
gpt-oss 20B
94%
Laguna S 2.1
85%
GLM-4.7 Flash
81%
Qwen3.6 35B
48%

Every case, every model

Mean of two seeds per case. Click any cell to open that case in the explorer.

0
100▢ = one seed only
ModelN1N2N3N4N5C2C3C4J1J2S1S2R1R2L1L5L3M1M2M3M4M5PL1PL2
Flash Next
Qwen3.8 27B
gpt-oss 120B
Gemma 4 26B
Gemma 4 31B
Gemma 4 12B
gpt-oss 20B
Qwen3.6 35B
Laguna S 2.1
GLM-4.7 Flash

Explorer

The saved answer for every model, per seed — prompts, scores, token counts, grader detail.

Agent tasks

Four multi-turn tool-using tasks, two seeds each, graded by hidden unit tests in the task repos.

A1 inventory · two bug reports

Hidden tests; an untouched repo scores 5/14.

ModelScoreTestsMax runNotes
Flash Next
100%
28/281 minall hidden tests, both seeds
Qwen3.8 27B
100%
28/282 minall hidden tests, both seeds
gpt-oss 120B
96%
27/281 min
Gemma 4 26B
100%
28/281 minall hidden tests, both seeds
Gemma 4 31B
93%
26/283 min
Gemma 4 12B
68%
19/2831 minhit 30-min limit
gpt-oss 20B
89%
25/283 min
Qwen3.6 35B
93%
26/281 min
Laguna S 2.1
100%
28/284 minall hidden tests, both seeds
GLM-4.7 Flash
86%
24/282 min

A2 events · cursor pagination

Hidden tests; an untouched repo scores 1/14.

ModelScoreTestsMax runNotes
Flash Next
100%
28/283 minall hidden tests, both seeds
Qwen3.8 27B
100%
28/285 minall hidden tests, both seeds
gpt-oss 120B
93%
26/2812 minhit 40-turn limit
Gemma 4 26B
93%
26/286 min
Gemma 4 31B
93%
26/283 min
Gemma 4 12B
93%
26/2820 min
gpt-oss 20B
50%
14/283 minrepo unchanged
Qwen3.6 35B
79%
22/289 min
Laguna S 2.1
50%
14/2834 minrepo unchangedhit 30-min limit
GLM-4.7 Flash
79%
22/2835 minhit 30-min limithit 40-turn limit

A3 billing · month-end crash and proration

Hidden tests; an untouched repo scores 1/14.

ModelScoreTestsMax runNotes
Flash Next
100%
28/282 minall hidden tests, both seeds
Qwen3.8 27B
100%
28/289 minall hidden tests, both seeds
gpt-oss 120B
100%
28/282 minall hidden tests, both seeds
Gemma 4 26B
93%
26/286 min
Gemma 4 31B
93%
26/287 min
Gemma 4 12B
96%
27/2832 minhit 30-min limit
gpt-oss 20B
82%
23/285 minhit 40-turn limit
Qwen3.6 35B
96%
27/283 min
Laguna S 2.1
100%
28/2826 minall hidden tests, both seeds
GLM-4.7 Flash
71%
20/2821 minhit 40-turn limit

A4 ratelimit · exact sliding-window limiter

Hidden tests; an untouched repo scores 0/1.

ModelScoreTestsMax runNotes
Flash Next
100%
32/326 minall hidden tests, both seeds
Qwen3.8 27B
91%
29/3211 min
gpt-oss 120B
75%
24/323 minload-sensitive test failed
Gemma 4 26B
75%
24/3234 minhit 30-min limitload-sensitive test failed
Gemma 4 31B
81%
26/328 min
Gemma 4 12B
69%
22/3244 minhit 30-min limitload-sensitive test failed
gpt-oss 20B
44%
14/176 minhit 40-turn limit
Qwen3.6 35B
97%
31/3220 minload-sensitive test failed
Laguna S 2.1
34%
11/1734 minrepo unchangedhit 30-min limitload-sensitive test failed
GLM-4.7 Flash
75%
24/3225 minhit 40-turn limitload-sensitive test failed

Thinking budget

Completion tokens across the single-turn answers, and what thinking cost. All ran at max_tokens 32,768; OpenCode's own per-model output limits listed last. "Score w/ thinking" is "—" for models whose records contain no reasoning.

ModelTotal tokensMedianMaxHit 32K>16K Visible thinkingScore w/ thinkingScore directOpenCode limit
Qwen3.6 35B
895,719
14,704.532,76813224867.4—32,000
Gemma 4 12B
553,162
10,503.032,768294888.8—16,384
GLM-4.7 Flash
542,137
8,512.532,7687104852.1—16,384
Gemma 4 26B
510,477
9,885.032,768154892.0—16,384
Qwen3.8 27B
263,606
3,733.019,833024894.6—32,000
Flash Next
248,552
3,924.524,812034895.3—32,000
gpt-oss 20B
240,338
3,516.032,611024889.1—16,384
Laguna S 2.1
233,792
703.032,768251668.066.132,000
gpt-oss 120B
183,174
2,583.020,543024892.4—16,384
Gemma 4 31B
54,108
803.08,969000—89.316,384

Speed

Clean single-workload measurements (cache-busted prefill, 512-token decode). Prefill columns are prompts sized by characters, not tokens — the x value differs per tokenizer.

ModelPP @2K chars tok/sPP @8K charsPP @32K charsDecode tok/s
Qwen3.6 35B2,9154,6694,496
162
gpt-oss 20B2,5614,2014,304
136
Gemma 4 26B2,0763,2323,047
95
gpt-oss 120B1,1981,9091,981
90
Flash Next1,2182,1772,460
75
GLM-4.7 Flash1,7252,7101,702
71
Laguna S 2.18131,1091,030
48
Qwen3.8 27B567714710
46
Gemma 4 12B1,0371,6541,581
39
Gemma 4 31B382567558
19

Quality vs speed

One point per finished local model: combined score against clean decode speed. The top-right corner is the sweet spot — high quality at high tokens per second. Cloud models never enter the charts.

50 100 150 200 20 40 60 80 100 Decode speed (tok/s) Combined score (%) Flash Next 96.0 Qwen3.8 27B 95.0 gpt-oss 120B 92.2 Gemma 4 26B 91.8 Gemma 4 31B 89.4 Gemma 4 12B 87.8 gpt-oss 20B 85.9 Qwen3.6 35B 70.8 Laguna S 2.1 67.4 GLM-4.7 Flash 55.8

x = each model's first clean speed.py decode run (512 tokens, cache-busted); y = combined over all cases and tasks. Every finished local model has a clean speed run.

Prefill speed vs context depth

20k 40k 60k 80k 100k 120k 140k 160k 180k 200k 1,000 2,000 3,000 4,000 5,000 Context depth (prompt tokens) Prefill speed (tokens/s) gpt-oss 20B 3,173.3 t/s Flash Next 2,077.5 t/s gpt-oss 120B 1,669.1 t/s Qwen3.6 35B 1,428.3 t/s Gemma 4 26B 1,073.2 t/s Gemma 4 12B 755.5 t/s Laguna S 2.1 667.0 t/s Qwen3.8 27B 398.5 t/s Gemma 4 31B 248.5 t/s GLM-4.7 Flash 186.8 t/s

Cache-busted prompts grown toward 10K–200K tokens, single-token prefill, repetitions averaged; the x-axis is measured prompt_tokens, so curves stay honest across tokenizers. Collected by the speed.py pp_ctx stage; models appear once swept. The sweep postdates the table's first-run figures, so values may come from a different session.

Decode speed vs context depth

20k 40k 60k 80k 100k 120k 140k 160k 180k 200k 50 100 150 200 Context depth (prompt tokens) Decode speed (tokens/s) gpt-oss 20B 73.9 t/s Qwen3.6 35B 60.0 t/s gpt-oss 120B 47.2 t/s Flash Next 45.6 t/s Gemma 4 26B 29.6 t/s Gemma 4 12B 27.9 t/s Qwen3.8 27B 22.2 t/s Laguna S 2.1 18.9 t/s GLM-4.7 Flash 18.8 t/s Gemma 4 31B 8.4 t/s

Cache-busted prompts grown toward 10K–200K tokens, 512-token decode, repetitions averaged; the x-axis is measured prompt_tokens, so curves stay honest across tokenizers. Collected by the speed.py tg_ctx stage; models appear once swept. The sweep postdates the table's first-run figures, so values may come from a different session.

Loaded memory vs the 107.5 GiB usable ceiling

Laguna S 2.1
72.9 GiB
Flash Next
69.0 GiB
gpt-oss 120B
59.0 GiB
GLM-4.7 Flash
29.6 GiB
Qwen3.6 35B
28.3 GiB
Qwen3.8 27B
28.1 GiB
Gemma 4 31B
26.0 GiB
Gemma 4 26B
20.7 GiB
Gemma 4 12B
12.2 GiB
gpt-oss 20B
10.9 GiB

The OpenCode resident model stays loaded; loading a benchmark model evicts others. After a run, reload the resident model (after-run.sh).

☁ Cloud models Codex CLI · same cases and agent tasks · detail tables here; in the league they are flagged ☁ as not comparable

Run through codex exec on the user's ChatGPT Plus plan. Different harness, tokenizer and reasoning accounting than the local oMLX runs — scores are indicative, not directly comparable.

ModelCombinedSingleAgenticStrict /48 Tool answersPeeked30-min timeoutsPlan usage
gpt-6-astra · medium100.0100.0100.048/48000
5-hour 99%
weekly 45%
gpt-6-sol · medium99.499.3100.044/48000
5-hour 16%
weekly 32%
gpt-5.6-terra · medium94.393.599.143/48000
5-hour 32%
weekly 35%
gpt-6-luna · medium84.581.9100.030/48000
5-hour 1%
weekly 30%

Cloud case matrix

ModelN1N2N3N4N5C2C3C4J1J2S1S2R1R2L1L5L3M1M2M3M4M5PL1PL2
gpt-6-astra · medium
gpt-6-sol · medium
gpt-5.6-terra · medium
gpt-6-luna · medium

Method and caveats

How it was scored

24 single-turn cases (Python, Java, SQL, trace debugging, reasoning, code reading) × 2 seeds (500, 501), plus 4 multi-turn tool-using agent tasks (A1–A4) × 2 seeds graded by hidden unit tests in the task repos. Partial credit throughout; combined = mean over all 24 cases and 4 tasks, each weighted equally after averaging seeds, so agent work is 4/28 of it. An untouched repo scores A1 5/14, A2 1/14, A3 1/14, A4 0/1.

Thinking on, 32K tokens

Thinking was on for every model except gemma-4-31B-it-oQ6e-mtp: its 48 single-turn rows and 8 agent rows went out with empty sampling and no reasoning (the X-1 finding), so its score is a thinking-off run and is disclosed as such. Per-model expectations now live in models.json: `request_extra` sends `enable_thinking` per request for the Gemma 4 models (whose chat template defaults thinking off), and `thinking_default` marks models that reason by server default. run.py merges `request_extra` into every request and fails the warmup when a model expected to think returns no reasoning. Sampling came from ~/.omlx/model_settings.json per model. Derived from the records: Gemma 4 31B returned no reasoning on any row — it ran with thinking off.

The gpt-oss Harmony bug

oMLX 0.7.0rc1 has a Harmony parser bug that silently drops completions whose first message is a tool call (empty stop). The harness resamples blank replies up to 16 times, but on the unpatched build gpt-oss-20b still lost whole agent runs to it. Agentic scores reflect oMLX + model, not the model alone.gpt-oss 20B’s agent runs are recorded as patched with ../patches/omlx-0.7.0rc1-harmony-tool-calls.patch; patch state unconfirmed — the installed bundle checks UNPATCHED; treat all agent runs as unpatched. The runs logged as unpatched for gpt-oss 20B needed 91 blank-reply resamples, 4 of 8 runs were blocked and agentic was 49.3%; the reruns recorded as patched logged 4, with 0 blocked, for 66.3%. Patch state is unconfirmed — treat every agent run as unpatched. The replaced runs stay in agent-results.jsonl and agent-transcripts-superseded/.

Load-sensitive tests

A4's test_amortized_constant_time and test_idle_keys_are_forgotten fail on a busy machine. a4_idle_regrade.py rebuilds saved A4 diffs and re-grades them on an idle machine (state/a4-idle-regrade.json); those rechecks are reported separately, and the tables keep the live grades. After the run, 16 A4 runs were rebuilt from their stored diffs and re-graded on the idle machine; every one reproduced its live score.

Wall clock and memory

Agent runs stop at 30 minutes or 40 turns; Laguna hit the clock on A2 and A4 both seeds. Loading a benchmark model evicts others on this 128 GiB M5 Max — Flash Next alone stays resident at ~69.5 GiB, and OpenCode traffic during a run contaminates speed numbers, so speed was measured in a dedicated clean pass.

Cloud models

The same cases and tasks ran through Codex CLI on a ChatGPT Plus plan. Cloud results are reported only in their own band: different harness, different tokenizer, reasoning tokens accounted differently, and every run spends real plan quota.

What v1 said, and why v2 withdrew it

v1 (thinking off, 2,000 tokens) recommended switching build/plan to Qwen3.6-35B. With thinking on and a 32K budget, 35B truncated 13 of 48 answers and fell to 70.8 — that recommendation is withdrawn. v1 and v2 scores are not comparable.