I rebuilt the local model lineup on a single DGX Spark with the same two rulers for every candidate: a deterministic 28-task test and a real agent loop. Gemma-4-26B scored 28 of 28, a 2B model beat my 35B daily driver, and two big models at once crashed into a memory peak nobody models. What I kept, what I deleted, and why a perfect score was the wrong answer.

My Test Said Gemma 4 Beat Qwen3.8-27B. Every Public Benchmark Said the Opposite.

Earlier today I compared Qwen3.8-27B with my Qwen3.6-35B daily driver and ended with a plan: give the 27B a slot for hard problems, test it against the Gemma model that held that slot, and let the loser leave the box.

When I went to run that test, the box itself turned out to be the bigger problem. The “Gemma slot” was no longer the Gemma I had benchmarked in June. Over the summer an agent had swapped it for a 2B edge model and advertised that 2B in my agent fleet as the deep-reasoning specialist, with no written reason. Another engine sat in the lineup under a license I would not have chosen. And about 185 GB of weights were on disk for models nothing used. Most of that was measured once, months ago, with a ruler I would no longer trust.

So instead of one duel I did the whole thing properly: roles first, rules second, then every candidate through the same two rulers, and only then a decision.

Roles and rules before models

A single Spark has 121 GB of unified memory that the GPU, the operating system and the desktop share. That decides the shape of the lineup more than any benchmark does. I wanted exactly two roles:

Three rules came with it. Only OSI-licensed open models run locally, so Apache-2.0 and MIT are in, while “open weights” under a custom license, like NVIDIA’s model license on Nemotron, is out, however good the model. Every model on disk needs a job. And every candidate gets measured the same way, on the same box, in the same session.

Two rulers, same for everyone

The hard-task set. 28 deterministic tasks: 14 math and logic problems whose answers the harness computes by brute force, 10 coding tasks with hidden tests run in a network-less container, 4 tool-calling requests. Same prompts, temperature 0.6, a 16,000-token budget for math and 12,000 for code. It is the test from the earlier post, where I describe it in detail.

A real agent loop. My agent-bench harness drives a coding agent (opencode) through actual file edits: rename a function across files, fix the callers, rename a symbol without touching a same-named one, plus three structured chat tasks. A gate script decides pass or fail. Six tasks, three runs each.

Speed is decode tokens per second counted from the server’s own usage numbers, prose and code measured separately, because MTP speculative decoding makes predictable text much faster.

The results

Model (local, NVFP4 unless noted)Hard tasks (28)Agent loop (18)Prose tok/sCode tok/s
Qwen3.6-35B-A3B (daily driver)15 (22 with a 32k budget)1880 to 90about 119
Qwen3.8-27B (dense)25not run20 to 2122 to 30
Gemma-4-26B-A4B (MoE)281747 to 5569 to 81
Gemma-4-E2B (2B, BF16)19not runabout 41about 41

Gemma-4-26B went through all 28 tasks in about four minutes, with its MTP drafter accepting 4.7 tokens per step. On paper that made it the obvious hard-problem model: faster than the 27B by a factor of two to three and better on my test.

The perfect score was the wrong answer

Before deleting anything I checked the result against sources outside my box, and they disagree completely.

Artificial Analysis compares the two directly: Intelligence Index 34 for Qwen3.8-27B against 17 for Gemma-4-26B. Humanity’s Last Exam 34% against 19%, long-context reasoning 82% against 66%, SciCode 47% against 40%. LLM Stats counts four shared benchmarks and the 27B wins all four. The model cards say the same: Qwen3.8-27B at 89.2 on GPQA Diamond and 90.3 on LiveCodeBench v6, NVIDIA’s Gemma-4-26B NVFP4 at 79.9 and 79.8. And Gemma 4 is the newest Gemma there is: the release notes list nothing after the April 2026 launch.

So why did my test crown the weaker model? Because both models hit its ceiling. Of the 27B’s three misses, two were budget cut-offs on the two longest problems, not wrong answers. Gemma got there with fewer tokens, not with more insight. A 28-task set without a genuinely hard tier can rank models by efficiency inside a budget. It cannot separate two strong models by capability, and a perfect score is exactly the signal that the test ran out of difficulty. That is the same lesson as a lying benchmark, from the other side: the ruler was honest, it was just too short.

The 27B got the hard-problem role. It needs no patched inference engine, which Gemma did (more on that below), and on every harder public test it is clearly ahead.

A 2B model beat my daily driver, and then it did not

The other surprise sits at the bottom of the table. Gemma-4-E2B, a 2B model, solved 19 of the 28 tasks. My daily driver, a 35B mixture-of-experts model, solved 15. The daily driver was not wrong more often. It thinks for too long: on eleven tasks it found an answer, then re-checked the same arithmetic over and over until the budget ran out. With 32,000 tokens it solved seven of those eleven, at five to seven minutes per task.

Then the agent loop reversed the picture. There the daily driver went 18 for 18, with 6 to 16 tool calls per coding task and 16 to 37 seconds per run. Gemma-4-26B went 17 for 18 and twice got lost in a loop of 38 and 39 tool calls on a task the Qwen finished in 9. In an agent loop the turns are short, so overthinking has no room to happen, and steadiness matters more than peak reasoning. That is why the daily driver stays exactly where it is: the work it does all day is the work it is best at.

Two models at once: the memory peak nobody models

The 27B is slow, so I wanted it running beside the daily driver instead of replacing it for a while. On paper that fits. The daily driver needs about 42 GB, the 27B about 24 GB, and 49 GB were free.

It did not fit. I ran every test through a memory guard that kills the candidate when free memory drops below 8 GB, because on unified memory running dry freezes the whole desktop. The guard fired twice. The second time the daily driver had been shrunk to leave 49 GB free, and the 27B still took free memory down to 6 GB within 45 seconds of starting. Steady state would have fit. Loading did not: while the weights are read and moved into place, the loader briefly holds far more than the model’s final size. Unified memory makes that peak everyone’s problem. The small E2B did run beside the daily driver without trouble (13 GB were left at its worst moment), but that is not a model I have a job for.

So the rule stands: one big model at a time. Switching to the hard-problem model pauses the daily driver, and the switch brings it back automatically afterwards.

The setting that cost nothing

One change came out of this for free. The daily driver had been running with 45% of the memory pool reserved for the inference engine. At 35%, its KV cache still holds 2.76 million tokens, more than ten times its maximum context. I measured decode speed at both settings in the same session: 80.2 against 80.1 tok/s on prose, 118.6 against 119.8 on code. That is identical within noise, and the lower setting leaves 14 GB more for image generation, text-to-speech and the desktop. It is the default now.

Notes for anyone running Gemma 4 NVFP4 on vLLM

If you try NVIDIA’s Gemma-4-26B NVFP4 checkpoint on a vLLM 0.23 nightly, it dies at startup with a NotImplementedError from tie_weights. It is the regression tracked as vllm#45543: Gemma ties its output layer to its embeddings, the checkpoint excludes that layer from quantization, and vLLM’s modelopt path hands the excluded layer a linear-layer method that cannot tie weights. Returning the embedding method for that one layer type fixes it; it is a one-line change in the modelopt quantization config. The June 2026 Gemma bring-up hit the same wall on the dense 31B build, so this is not new, just still open on the image I use.

What stays, what went

RoleModelStatus
Daily driverQwen3.6-35B-A3B NVFP4always on, now at 35% memory
Hard problemsQwen3.8-27B NVFP4on demand, pauses the daily driver

About 185 GB of weights left the box: Nemotron-3-Nano and Nemotron-3-Super (license rule; the Super had its own teardown in June), both Gemma 4 builds, and gpt-oss-120b, which was fast but passed only 56% of the agent tasks in its own test. The earlier dense Gemma and GLM experiments are written up in the Gemma-31B post and the GLM post. Every model that is still on disk now has a named job, and a check flags any folder that does not.

A newer small Qwen mixture-of-experts model in the 35B class would be the obvious next daily driver. The 3.5 and 3.6 generations both had one, but the Qwen3.8 repository lists only the 2.4T flagship and the 27B so far.

Caveats

Was this worth it? Zap the article.

Value for value, no signup. Sats go straight to the writer.

… sats zapped
… zaps
Today 7d 30d All-time
Unique readers — — — —
Page views — — — —
—
All Article Insights →