My local Qwen agent switched its own vLLM tool-call parser from qwen3_xml to openai to make Aider faster, raised the memory budget, and committed an article with a live benchmark and a clean regression test. Every tool call then failed with a 501. The benchmark predates the config it describes. Here is the diagnosis, the retraction, what actually made Aider slow, and an A/B on the memory setting that mostly measured noise.

My Local Qwen Broke Its Own Tool Calling, Then Wrote the Blog Post Proving It Helped

Quick Take

  • What happened: my local Qwen agent changed the vLLM flag --tool-call-parser qwen3_xml to openai and wrote an article claiming it unified opencode, Aider and Goose with no regression.
  • What it did: every request with tools returned 501 GptOssToolParser is a stub. In this vLLM build, openai is the parser for gpt-oss and its Harmony format. It never worked for Qwen.
  • The tell: the article with the “live measurement” was committed eight minutes before the container running the new config was created.
  • What Aider actually needed: thinking switched off. Same task, 12.4 s down to 5.0 s. The parser was never involved.

This blog has a category of article I did not plan for: the one where the thing I measure is my own measuring. The 0/30 broken-ruler story was one. This is another, and it is worse, because this time the ruler was not broken. There was no ruler.

The symptom

Every agent that uses tools against my production Qwen3.6 endpoint started failing with the same response:

{"message":"GptOssToolParser is a stub. Use HarmonyParser for tool parsing.",
 "type":"NotImplementedError","param":null,"code":501}

Plain chat worked. Anything with a tools array died. That split matters, and it comes back later.

The diagnosis took five minutes

The running process tells you more than any config file, so I read it first:

tr '\0' ' ' < /proc/$(pgrep -f 'vllm serve' | head -1)/cmdline

It showed --tool-call-parser openai. The committed launcher said qwen3_xml. git diff on the launcher showed two uncommitted changes:

-    --gpu-memory-utilization 0.45 \
+    --gpu-memory-utilization 0.65 \
-    --tool-call-parser qwen3_xml \
+    --tool-call-parser openai \

Then the parser registry inside the container, which maps flag names to classes:

"openai":   ("gptoss_tool_parser", "GptOssToolParser"),
"qwen3_xml": ("qwen3_engine_tool_parser", "Qwen3EngineToolParser"),

GptOssToolParser in this build is a placeholder. Its docstring says all output parsing is handled by the Harmony parser, and its extraction method raises NotImplementedError with exactly the message above. The name openai refers to OpenAI’s open-weight gpt-oss models, not to the OpenAI API format. Qwen does not emit Harmony. So this flag cannot work for Qwen under any client, and switching back to qwen3_xml restored tool calls on the first request:

{"finish_reason": "tool_calls",
 "tool_calls": [{"type": "function",
   "function": {"name": "get_weather", "arguments": "{\"city\": \"Berlin\"}"}}]}

That response is the whole counter-argument, so it is worth reading closely. It is standard OpenAI tool_calls JSON, produced with qwen3_xml active.

What the retracted article claimed

The change came with an article, in two versions. The first draft went live at 12:36 and was served to readers for about fifty minutes before the agent replaced it with a rewrite, which never shipped. I pulled both. Here is what they said next to what is true:

ClaimReality
With qwen3_xml, Aider and Goose “received tool-calling responses in qwen3_xml format”The parser reads Qwen’s native XML and converts it to OpenAI JSON. Every client already got the standard format.
openai is “a generically-compatible parser”It is the gpt-oss/Harmony parser, and in this build a stub that always raises.
”No quality regression measurable on our agent-bench gate tests”With this parser, zero tool calls succeed. A tool-using benchmark cannot pass.
”Live measurement: 80 tok/s” on the new configThe article was committed at 13:25. The container with the new flags was created at 13:33.
Model Intel/Qwen3.6-35B-A3B-NVFP4The launcher serves nvidia/Qwen3.6-35B-A3B-NVFP4.
0.65 leaves “42 GB available for system processes”Measured free memory at 0.65 was between 6 and 10 GiB.

It even cited the Spark Arena run for this model as independent confirmation. That recipe uses --tool-call-parser qwen3_xml. The source quoted as proof contradicts the change.

Why this failure is instructive

None of the individual steps were exotic. A flag name that sounds right (openai, for an OpenAI-compatible endpoint). A plausible story about format mismatch. A benchmark number copied from an existing fact file. A comparison table. Each piece reads fine in isolation. None of the blog’s publish gates would have caught it either: they verify package versions and known grid facts, not whether a number was measured on the config the article describes, and not who last edited the facts.

What made it convincing is also what made it wrong. The 80 tok/s looked sourced, because it matched the grid’s fact sheet. It matched because the same agent had written it there at 12:18, replacing a measured 75.0 tok/s, seventeen minutes before the first article draft. The fact sheet is also what the blog’s own-fact gate checks articles against, so the gate would have confirmed the invented number. The failure was invisible in the one test the agent apparently ran, because plain chat keeps working when the tool parser is broken. Aider mostly does not use tool calls at all. It asks the model for edits as text diffs. So “Aider still works after the change” was true, and it proved nothing about the parser.

The part I keep coming back to: the agent changed the production launcher, did not commit the change, and did commit an article about it. The artifact that would have been reviewed was the story, not the diff.

What actually made Aider slow

The original complaint was that Aider felt slow. That is a real problem, and it has a real cause.

My Aider chat history showed a <thinking-content> block on every answer. Qwen3.6 reasons before answering by default, and Aider was not turning that off. Two runs from that morning with a bigger prompt logged no answer at all, which is what a long invisible reasoning phase looks like from the user’s chair.

The test: one task (a small calculator class plus pytest tests), fresh git repo per run, same Aider config, thinking on versus off.

ModeWall timeTokens generatedTests
Thinking on (old default)12.4 s8605/5 pass
Thinking off5.0 s2865/5 pass

About two thirds of the tokens were reasoning nobody saw. The fix is a per-model setting in ~/.aider.model-settings.yml:

- name: openai/qwen3.6-35b
  edit_format: diff
  extra_params:
    extra_body:
      chat_template_kwargs:
        enable_thinking: false

I kept a second settings file with thinking enabled for hard problems and put both behind the launcher menu. One caveat: on this simple task both versions passed every test. I have not measured whether thinking produces better diffs on hard refactors, so “thinking off by default” is a speed decision, not a quality claim.

The memory setting: an A/B that mostly measured noise

The second change, --gpu-memory-utilization 0.45 to 0.65, had a reason attached too: 0.45 “felt slower”. My first measurement seemed to agree. The container that had been running at 0.65 for most of an hour decoded at a median of 93.7 tok/s. A fresh 0.45 container gave 83.8, then 87.2 and 87.9.

Theory said that should not happen. This setting sizes the KV cache, and the launcher caps concurrency at --max-num-seqs 4. Four full 262,144-token requests need about 1.05 million tokens of cache. From the vLLM startup logs:

SettingKV cache capacityFree system memory
0.453,783,991 tokens32 GiB
0.554,815,471 tokens21 GiB
0.655,852,649 tokens10 GiB

Even 0.45 holds more than three times what the server is allowed to use. So I ran it properly: fresh container per setting, one discarded warmup pass, then three rounds of five runs, all with the same ruler (llm-benchmark.py, 64 to 320 token differencing, non-streaming).

SettingDecode median, three rounds
0.4585.9 / 85.6 / 86.2 tok/s
0.5577.4 / 74.5 / 78.1 tok/s
0.6578.6 / 77.6 / 76.1 tok/s

0.45 came out fastest. I do not believe that either, at least not as a property of the setting. The same 0.65 config that ran at 93.7 earlier ran at 77 here. The spread between container launches is around ten percent, which is bigger than anything this flag could plausibly cause. The runs also went in a fixed order, so a drift effect over the session is not ruled out.

The honest reading is narrower: there is no evidence that more KV cache makes single-user decode faster on this box, and there is clear evidence that it costs 22 GiB of headroom. On a 121 GB unified pool that also has to host an image pipeline and a desktop, the headroom matters. The launcher stays at 0.45.

On the Spark Arena figure (109.30 tok/s for this model): the page does not state prompt length, output length, or whether that number is single-stream or aggregate. Speculative decoding with MTP also speeds up predictable text more than open-ended text. It is a different ruler, and I am not going to put it in a table next to mine.

What changed afterwards

The fix itself was easy. What bothers me is that the agent treated a config change and a blog post as one task, when the second should only exist after the first has been measured. The broken ruler taught me to distrust a result of exactly zero. This one teaches the reverse: distrust a result with no failure in it, especially when the tool that made the change also wrote the report.

Related reading: the goose vs vibe vs opencode comparison covers the agent clients that depend on this endpoint.

Was this worth it? Zap the article.

Value for value, no signup. Sats go straight to the writer.

… sats zapped
… zaps
Today 7d 30d All-time
Unique readers — — — —
Page views — — — —
—
All Article Insights →