I checked my DGX Spark's Qwen3.6-35B against its Spark Arena recipe, found a benchmark tool that counts stream events instead of tokens, then ran a harder deterministic A/B against Qwen3.8-27B: 24 of 28 against 15, mostly because the 35B overthinks. Along the way the benchmark crashed production three times and exposed a watchdog bug.

Qwen3.8-27B Beat My Qwen3.6-35B on a Harder Test. The Test Also Crashed My Server.

My daily driver on a single DGX Spark is nvidia/Qwen3.6-35B-A3B-NVFP4 on vLLM. Two questions came up on the same afternoon. Is it still set up as well as the box allows? And now that a newer generation of Qwen is out, is there a stronger model that fits on one Spark and is worth the switch?

Answering the first one turned up a benchmark tool that under-counted my speed by a factor of three. Answering the second one needed a harder test than the one I had, because my old test could no longer tell good models apart.

Checking the setup against the recipe

The reference is a Spark Arena run of the same NVFP4 checkpoint at 109.3 tokens per second. I put its launch flags next to my production launcher, line by line. They match: FP8 KV cache, FlashInfer attention, the Marlin MoE backend with VLLM_MARLIN_USE_ATOMIC_ADD=1, and MTP speculative decoding with three draft tokens on a Triton MoE path.

One flag differs. The recipe runs --gpu-memory-utilization 0.65, mine runs 0.45. On the Spark that fraction is taken from the whole unified memory pool that the operating system, the desktop and every other container also live in. At 0.45 vLLM already reports a KV cache of 3,753,598 tokens, which is 14 times the model’s full 262k context. With --max-num-seqs 4 there is no way to use more. Raising it to 0.65 would move roughly 24 GB from the rest of the machine into a cache that stays empty, and it does not touch single-stream speed. I kept 0.45.

So the configuration is the recipe. The next question was whether the speed is the recipe’s speed too.

The ruler that counted packets

I ran llama-benchy, the tool behind the Spark Arena numbers, against my server. It reported about 31 tokens per second for generation. My own benchmark had measured 79 the same hour. A factor of 2.5 between two rulers is not noise, so I read the tool’s client code. The loop that consumes the stream does this for every server-sent event that carries text:

result.total_tokens += 1
result.token_timestamps.append(chunk_time)

It counts events, not tokens. Without speculative decoding that is the same thing, because the server sends one token per event. With MTP it is not: each decode step can accept several draft tokens, and vLLM streams them together in one event. On my server an event carried 2.3 tokens on prose and 3.4 on code. Thirty-one events per second times roughly 2.5 tokens per event is the 78 that my own tool measured.

To get a number I trust, I count what the server says it generated. Every OpenAI-compatible response carries usage.completion_tokens, and the decode rate is that count divided by the time between the first and the last streamed token. Measured that way, over four runs of 512 tokens each:

WorkloadDecode, medianRange
Code (an LRU cache module)104.3 tok/s101.4 to 106.0
Prose (continue a story)69.7 tok/s61.4 to 73.6

On code the setup reaches the recipe’s number. On prose it does not, and the reason is the same one that broke the ruler: MTP’s draft head guesses predictable text better. Code is full of predictable tokens, a story is not. I have seen the same effect before with EAGLE, where one run swung between 14 and 31 tokens per second depending on the content. A single tok/s figure for a speculative setup is only half a statement unless it says what text was generated.

The dashboard’s speed test on my box now measures both workloads and prints them side by side, so the number I look at every day can be compared with the recipe without a footnote.

This is not the first time a ruler lied to me on this box. A harness that hung made two capable models score exactly zero, and an earlier leaderboard comparison came out at 239 against 71. The working rule from the measurement-traps post held again: when two rulers disagree by a round factor, read the ruler’s code before believing either number.

What is new for a single Spark

The second question needed a look at what has shipped since I settled on the 35B. The candidates that can run on one Spark, from public reports that I have not reproduced myself:

The 27B was the only candidate cheap enough to test properly, so I tested it.

A test that can still tell models apart

My existing reasoning probe was useless for this. The 35B already scores 7 of 7 on its hard set and 10 of 10 on the easy one, so a better model cannot show up as better. I wrote a harder set with three parts, all scored without a judge model:

Before any model sees a task, the harness runs a self-check: every hidden test must pass against a reference implementation, and the answer checker must accept a correct answer and reject a wrong one. That self-check caught two bugs in my own harness on the first run. The seating puzzle I had written had no solution at all, and the guard that stops a model from solving the calculator task with eval() flagged the test file itself, because the test contained the string it was looking for. Both would have cost a model points it had earned. Two more negative controls prove the tests bite: a solution that uses eval() and a merge function that returns its input unchanged both fail.

To keep my production model serving while the test ran, the 27B ran on OpenRouter’s free endpoint for Qwen3.8-27B and the 35B ran on my Spark. That is not a perfectly fair pairing: the hosted model most likely runs at higher precision than the NVFP4 build I would deploy. NVIDIA’s own model card puts the gap at about one point on GPQA Diamond (88.9 against 88.0), so the hosted run is a slightly optimistic stand-in. Every task used the same prompt, temperature 0.6, top-p 0.95, and a 16,000-token budget for math.

Results

Same 28 tasks, same prompts, same budgets, temperature 0.6, one sample each:

Math (14)Code (10)Tools (4)TotalTokens spent
Qwen3.8-27B (hosted)128424 / 28123,564
Qwen3.6-35B-A3B NVFP4 (local)57315 / 28244,338

The gap is large, and the reason for it matters more than the score. Of the 35B’s 13 failures, 11 were not wrong answers. The model never answered: it spent the whole budget, 16,000 tokens on math and 12,000 on code, and was cut off inside its reasoning. Only one of its math answers was actually wrong (871 where the answer is 581). The 27B ran out of budget four times.

I read one of the cut-off traces in full. On the “last three digits of 7^2026” task the 35B found 649 early, verified the cycle length, then kept verifying: the same multiplication chain again, the exponent reduction again, the square of 343 by hand, each block ending in “Correct.” or “Wait,”. It was right the whole time and never stopped checking. The 27B solved the same task in 1,267 tokens.

Median tokens for a solved task: 1,239 for the 27B, 3,408 for the 35B. The 35B needs almost three times as much reasoning to get to an answer, and on a third of this set it does not get there within 16k.

To check that the budget, not the model, was the limit, I re-ran the 35B’s cut-off tasks with 32,000 tokens. It solved 7 of the 11: the tiling count, the powers of 7, the knight, the coins, the permutations, the probability and the island counter, each after 16,000 to 31,000 tokens. That lifts the 35B to 22 of 28, still behind the 27B’s 24, which it scored on half the budget. Two coding tasks, the expression evaluator and the LRU cache, stayed unsolved even at 32,000 tokens: the model reasoned for the entire budget and never wrote the code block. So the budget explains most of the gap, not all of it, and the price of closing it is paid in time. At my measured decode rates a 25,000-token answer is five to six minutes of waiting.

The single tool-calling miss of the 35B deserves an asterisk. In the scored run its first call was exactly right (search_files, pattern TODO, path /src, glob *.py), but the harness requires exactly one call and it emitted more than one. In four re-runs of the same request it made a single correct call every time. I count it as a miss because that is what the run produced, but it is noise, not a capability gap.

What the public rankings say

Twenty-eight tasks is a small sample, so I checked the result against two larger sources.

Artificial Analysis compares the two models directly. Their Intelligence Index puts the 27B at 34 and the 35B-A3B at 18. The gap is widest on agentic work (48% against 5% on their AutomationBench) and still clear on Humanity’s Last Exam (34% against 22%) and SciCode (47% against 37%). Their hosted speed measurement has the opposite sign: 43 tokens per second for the 27B, 125 for the 35B. One detail disagrees with my run. They test the 27B at its highest reasoning effort (xhigh), where it spends more output tokens per task than the 35B, 67k against 35k. At the default effort I used, it spent fewer. Qwen3.8 exposes that effort as a per-request setting, a dial the 35B does not have.

Qwen’s own model cards point the same way. The Qwen3.8-27B card compares it with the older Qwen3.6-27B rather than with my 35B, but the Qwen3.6-35B-A3B card reports several of the same benchmarks. Put side by side: SWE-bench Pro 61.7 against 49.5, LiveCodeBench v6 90.3 against 80.4, GPQA Diamond 89.2 against 86.0, HLE 30.8 against 21.4. These come from two separate releases with their own harnesses, so treat the pairs as direction, not as exact margins. The direction is the same everywhere: clearly smarter, roughly three times slower.

A smaller Qwen3.8 MoE in the 35B-A3B class would be the obvious successor for this box. The 3.5 and 3.6 generations both had one. As of writing, the Qwen3.8 repository lists only the 2.4T-A95B flagship and the 27B, and I found no announcement of a smaller MoE.

The benchmark crashed my production model, three times

This part was not planned. Qwen’s model card recommends a different sampling preset for “thinking, general” tasks than the one I used: temperature 1.0, top_k 20 and presence_penalty 1.5, a penalty meant exactly against the kind of repetitive re-checking I had just watched. To be fair to the 35B I re-ran it with that preset.

Eight minutes in, the server died. The kernel log had one line I had never seen on this box in the two weeks of journal I keep:

NVRM: Xid (PCI:000f:01:00): 31, ... name=VLLM::EngineCor ... MMU Fault: ... Fault is of type FAULT_PDE

vLLM’s own trace ended in CUDA error: an illegal memory access was encountered, raised in the FlashInfer attention backend while building metadata for the next decode step. Docker restarted the container and the model was back in about three minutes.

The obvious suspect was the one thing I had changed, the penalty. I re-ran the preset once more to test that. The engine died again within 20 seconds, same Xid, same trace. Two for two, and I wrote it down as a finding: sampling penalties crash this stack. There is even an upstream CVE for penalties crashing vLLM’s speculative decoding, for a different proposer and fixed versions ago, which made the story feel complete.

The third crash took the story apart. It happened on the 32k re-run, with no penalty at all, temperature 0.6, the exact settings that had run clean for 18 minutes earlier. The scheduler dump that vLLM writes on a fatal error showed what all three crashes had in common: four requests running at once, every slot of --max-num-seqs 4 taken, each about 2,000 tokens into a long generation. The penalty had been a coincidence. What I actually have is an intermittent fault in long, fully parallel decoding on this particular combination of MTP speculation, asynchronous scheduling and Qwen3.6’s hybrid attention layers. In normal use my agents rarely run four long requests at once, which is why it had never shown up before. A first data point: the last 32k re-run, at two requests in parallel instead of four, ran 25 minutes without a fault. That is not proof, but it fits. Isolating it is the next job, on a quiet evening, one flag at a time.

There was a second failure mixed into the same afternoon, and this one was mine. Between the crashes the model was also restarted once by my own watchdog. It probes the server every five minutes with a one-token request and restarts the engine after two failed probes, a rule written after the engine once hung for three hours while its health endpoint kept answering. With four long requests holding every slot, the probe sat in the queue, timed out twice, and the watchdog killed a perfectly healthy engine in the middle of its work. The fix is small: when the probe fails, the watchdog now reads vLLM’s generation_tokens_total counter twice, 20 seconds apart. If it moves, the engine is busy, not hung. The three-hour hang would still be caught, because in that case the counter stood still. The watchdog’s test suite had also quietly rotted since its last rewrite, with 9 of 26 checks failing against the current script. The new one has 23 checks, and against the old script exactly the four “busy” checks fail, which is what they are for.

The 27B on my own Spark

Everything above used the hosted 27B. The question that decides whether it earns a place on the box is how it runs there, so I pulled the NVIDIA NVFP4 checkpoint and served it with the same recipe as production minus the MoE flags: FP8 KV cache, FlashInfer attention, MTP with three draft tokens, 0.45 of memory. Production paused for the test, because the two do not fit side by side. At 0.45 vLLM gave the 27B a KV cache of 1,215,528 tokens, about a third of what the MoE gets, because a dense model’s cache per token is larger. The first start took nine and a half minutes, most of it kernel compilation.

Speed, with the same usage-counted ruler as before:

WorkloadQwen3.8-27B NVFP4Qwen3.6-35B-A3B NVFP4
Prose20 to 21 tok/s70 to 99 tok/s
Code22 to 30 tok/sabout 104 tok/s

That is the third-party number, confirmed: three and a half to four and a half times slower. MTP helps, with 2.2 to 2.5 accepted tokens per step, but a dense 27B has to stream all of its weights through memory for every step, where the MoE touches about 3B. On a box with 273 GB/s of memory bandwidth that difference is the whole story.

Quality held up under quantization. The local NVFP4 build solved 25 of 28, one more than the hosted run (one extra coding task, within noise), with a median of 1,035 tokens per solved task. The full set took 26 minutes at three requests in parallel.

Put speed and verbosity together and the picture is less lopsided than the tok/s column suggests. A median solved task costs the 27B about 1,035 tokens at 21 tok/s, roughly 50 seconds. The 35B needs about 3,400 tokens at 80, roughly 45 seconds. On easy work they finish at about the same time. On hard work the 27B finishes, and the 35B either spends five to seven minutes or does not answer.

What I am keeping

The 35B stays the daily driver. Most of what my agents send it is short, interactive and tool-heavy, and there four times the decode speed wins.

The 27B gets a slot in the model switcher for the hard problems, the ones where the 35B thinks itself into the budget. That is the job my Gemma slot was set up for, so the next test is the same 28 tasks against Gemma, and the loser leaves the box.

The afternoon also left three fixes that matter more than the scoreboard: a benchmark ruler I no longer trust without reading its code, a watchdog that now tells busy from hung, and a crash under full parallel load that I can now reproduce and will isolate next.

Caveats

Update, same evening: the Gemma test happened, with a twist. Gemma-4-26B scored 28 of 28 on this set, and every public benchmark still ranks it far below the 27B, because the set ran out of difficulty. The 27B got the hard-problem role; the full lineup test is in My Test Said Gemma 4 Beat Qwen3.8-27B.

Was this worth it? Zap the article.

Value for value, no signup. Sats go straight to the writer.

… sats zapped
… zaps
Today 7d 30d All-time
Unique readers — — — —
Page views — — — —
—
All Article Insights →