My Test Said Gemma 4 Beat Qwen3.8-27B. Every Public Benchmark Said the Opposite.
Earlier today I compared Qwen3.8-27B with my Qwen3.6-35B daily driver and ended with a plan: give the 27B a slot for hard problems, test it against the Gemma model that held that slot, and let the loser leave the box.
When I went to run that test, the box itself turned out to be the bigger problem. The “Gemma slot” was no longer the Gemma I had benchmarked in June. Over the summer an agent had swapped it for a 2B edge model and advertised that 2B in my agent fleet as the deep-reasoning specialist, with no written reason. Another engine sat in the lineup under a license I would not have chosen. And about 185 GB of weights were on disk for models nothing used. Most of that was measured once, months ago, with a ruler I would no longer trust.
So instead of one duel I did the whole thing properly: roles first, rules second, then every candidate through the same two rulers, and only then a decision.
Roles and rules before models
A single Spark has 121 GB of unified memory that the GPU, the operating system and the desktop share. That decides the shape of the lineup more than any benchmark does. I wanted exactly two roles:
- A daily driver, always on: agents, chat, vision, long context, speed.
- A hard-problem model, on demand: quality over speed, for the tasks where the daily driver gives up.
Three rules came with it. Only OSI-licensed open models run locally, so Apache-2.0 and MIT are in, while “open weights” under a custom license, like NVIDIA’s model license on Nemotron, is out, however good the model. Every model on disk needs a job. And every candidate gets measured the same way, on the same box, in the same session.
Two rulers, same for everyone
The hard-task set. 28 deterministic tasks: 14 math and logic problems whose answers the harness computes by brute force, 10 coding tasks with hidden tests run in a network-less container, 4 tool-calling requests. Same prompts, temperature 0.6, a 16,000-token budget for math and 12,000 for code. It is the test from the earlier post, where I describe it in detail.
A real agent loop. My agent-bench harness drives a coding agent (opencode) through actual file edits: rename a function across files, fix the callers, rename a symbol without touching a same-named one, plus three structured chat tasks. A gate script decides pass or fail. Six tasks, three runs each.
Speed is decode tokens per second counted from the server’s own usage numbers, prose and code measured separately, because MTP speculative decoding makes predictable text much faster.
The results
| Model (local, NVFP4 unless noted) | Hard tasks (28) | Agent loop (18) | Prose tok/s | Code tok/s |
|---|---|---|---|---|
| Qwen3.6-35B-A3B (daily driver) | 15 (22 with a 32k budget) | 18 | 80 to 90 | about 119 |
| Qwen3.8-27B (dense) | 25 | not run | 20 to 21 | 22 to 30 |
| Gemma-4-26B-A4B (MoE) | 28 | 17 | 47 to 55 | 69 to 81 |
| Gemma-4-E2B (2B, BF16) | 19 | not run | about 41 | about 41 |
Gemma-4-26B went through all 28 tasks in about four minutes, with its MTP drafter accepting 4.7 tokens per step. On paper that made it the obvious hard-problem model: faster than the 27B by a factor of two to three and better on my test.
The perfect score was the wrong answer
Before deleting anything I checked the result against sources outside my box, and they disagree completely.
Artificial Analysis compares the two directly: Intelligence Index 34 for Qwen3.8-27B against 17 for Gemma-4-26B. Humanity’s Last Exam 34% against 19%, long-context reasoning 82% against 66%, SciCode 47% against 40%. LLM Stats counts four shared benchmarks and the 27B wins all four. The model cards say the same: Qwen3.8-27B at 89.2 on GPQA Diamond and 90.3 on LiveCodeBench v6, NVIDIA’s Gemma-4-26B NVFP4 at 79.9 and 79.8. And Gemma 4 is the newest Gemma there is: the release notes list nothing after the April 2026 launch.
So why did my test crown the weaker model? Because both models hit its ceiling. Of the 27B’s three misses, two were budget cut-offs on the two longest problems, not wrong answers. Gemma got there with fewer tokens, not with more insight. A 28-task set without a genuinely hard tier can rank models by efficiency inside a budget. It cannot separate two strong models by capability, and a perfect score is exactly the signal that the test ran out of difficulty. That is the same lesson as a lying benchmark, from the other side: the ruler was honest, it was just too short.
The 27B got the hard-problem role. It needs no patched inference engine, which Gemma did (more on that below), and on every harder public test it is clearly ahead.
A 2B model beat my daily driver, and then it did not
The other surprise sits at the bottom of the table. Gemma-4-E2B, a 2B model, solved 19 of the 28 tasks. My daily driver, a 35B mixture-of-experts model, solved 15. The daily driver was not wrong more often. It thinks for too long: on eleven tasks it found an answer, then re-checked the same arithmetic over and over until the budget ran out. With 32,000 tokens it solved seven of those eleven, at five to seven minutes per task.
Then the agent loop reversed the picture. There the daily driver went 18 for 18, with 6 to 16 tool calls per coding task and 16 to 37 seconds per run. Gemma-4-26B went 17 for 18 and twice got lost in a loop of 38 and 39 tool calls on a task the Qwen finished in 9. In an agent loop the turns are short, so overthinking has no room to happen, and steadiness matters more than peak reasoning. That is why the daily driver stays exactly where it is: the work it does all day is the work it is best at.
Two models at once: the memory peak nobody models
The 27B is slow, so I wanted it running beside the daily driver instead of replacing it for a while. On paper that fits. The daily driver needs about 42 GB, the 27B about 24 GB, and 49 GB were free.
It did not fit. I ran every test through a memory guard that kills the candidate when free memory drops below 8 GB, because on unified memory running dry freezes the whole desktop. The guard fired twice. The second time the daily driver had been shrunk to leave 49 GB free, and the 27B still took free memory down to 6 GB within 45 seconds of starting. Steady state would have fit. Loading did not: while the weights are read and moved into place, the loader briefly holds far more than the model’s final size. Unified memory makes that peak everyone’s problem. The small E2B did run beside the daily driver without trouble (13 GB were left at its worst moment), but that is not a model I have a job for.
So the rule stands: one big model at a time. Switching to the hard-problem model pauses the daily driver, and the switch brings it back automatically afterwards.
The setting that cost nothing
One change came out of this for free. The daily driver had been running with 45% of the memory pool reserved for the inference engine. At 35%, its KV cache still holds 2.76 million tokens, more than ten times its maximum context. I measured decode speed at both settings in the same session: 80.2 against 80.1 tok/s on prose, 118.6 against 119.8 on code. That is identical within noise, and the lower setting leaves 14 GB more for image generation, text-to-speech and the desktop. It is the default now.
Notes for anyone running Gemma 4 NVFP4 on vLLM
If you try NVIDIA’s Gemma-4-26B NVFP4 checkpoint on a vLLM 0.23 nightly, it dies at startup with a NotImplementedError from tie_weights. It is the regression tracked as vllm#45543: Gemma ties its output layer to its embeddings, the checkpoint excludes that layer from quantization, and vLLM’s modelopt path hands the excluded layer a linear-layer method that cannot tie weights. Returning the embedding method for that one layer type fixes it; it is a one-line change in the modelopt quantization config. The June 2026 Gemma bring-up hit the same wall on the dense 31B build, so this is not new, just still open on the image I use.
What stays, what went
| Role | Model | Status |
|---|---|---|
| Daily driver | Qwen3.6-35B-A3B NVFP4 | always on, now at 35% memory |
| Hard problems | Qwen3.8-27B NVFP4 | on demand, pauses the daily driver |
About 185 GB of weights left the box: Nemotron-3-Nano and Nemotron-3-Super (license rule; the Super had its own teardown in June), both Gemma 4 builds, and gpt-oss-120b, which was fast but passed only 56% of the agent tasks in its own test. The earlier dense Gemma and GLM experiments are written up in the Gemma-31B post and the GLM post. Every model that is still on disk now has a named job, and a check flags any folder that does not.
A newer small Qwen mixture-of-experts model in the 35B class would be the obvious next daily driver. The 3.5 and 3.6 generations both had one, but the Qwen3.8 repository lists only the 2.4T flagship and the 27B so far.
Caveats
- One sample per task in the hard set, three runs per task in the agent loop. Differences of one or two passes are noise.
- The hard set saturates for strong models, which is the main finding of this post, not a footnote. The next model decision waits for a harder tier.
- The 27B was not run through the agent loop; its role does not involve interactive agent work.
- Public benchmark numbers come from the linked sources and use their own settings (Artificial Analysis runs the 27B at its highest reasoning effort). My own numbers are single-stream decode on one DGX Spark.
- The E2B and 27B speeds beside the daily driver are not reported, because the 27B never finished loading there.