My Local Qwen Broke Its Own Tool Calling, Then Wrote the Blog Post Proving It Helped
Quick Take
- What happened: my local Qwen agent changed the vLLM flag
--tool-call-parser qwen3_xmltoopenaiand wrote an article claiming it unified opencode, Aider and Goose with no regression.- What it did: every request with tools returned
501 GptOssToolParser is a stub. In this vLLM build,openaiis the parser for gpt-oss and its Harmony format. It never worked for Qwen.- The tell: the article with the “live measurement” was committed eight minutes before the container running the new config was created.
- What Aider actually needed: thinking switched off. Same task, 12.4 s down to 5.0 s. The parser was never involved.
This blog has a category of article I did not plan for: the one where the thing I measure is my own measuring. The 0/30 broken-ruler story was one. This is another, and it is worse, because this time the ruler was not broken. There was no ruler.
The symptom
Every agent that uses tools against my production Qwen3.6 endpoint started failing with the same response:
{"message":"GptOssToolParser is a stub. Use HarmonyParser for tool parsing.",
"type":"NotImplementedError","param":null,"code":501}
Plain chat worked. Anything with a tools array died. That split matters, and it comes back later.
The diagnosis took five minutes
The running process tells you more than any config file, so I read it first:
tr '\0' ' ' < /proc/$(pgrep -f 'vllm serve' | head -1)/cmdline
It showed --tool-call-parser openai. The committed launcher said qwen3_xml. git diff on the launcher showed two uncommitted changes:
- --gpu-memory-utilization 0.45 \
+ --gpu-memory-utilization 0.65 \
- --tool-call-parser qwen3_xml \
+ --tool-call-parser openai \
Then the parser registry inside the container, which maps flag names to classes:
"openai": ("gptoss_tool_parser", "GptOssToolParser"),
"qwen3_xml": ("qwen3_engine_tool_parser", "Qwen3EngineToolParser"),
GptOssToolParser in this build is a placeholder. Its docstring says all output parsing is handled by the Harmony parser, and its extraction method raises NotImplementedError with exactly the message above. The name openai refers to OpenAI’s open-weight gpt-oss models, not to the OpenAI API format. Qwen does not emit Harmony. So this flag cannot work for Qwen under any client, and switching back to qwen3_xml restored tool calls on the first request:
{"finish_reason": "tool_calls",
"tool_calls": [{"type": "function",
"function": {"name": "get_weather", "arguments": "{\"city\": \"Berlin\"}"}}]}
That response is the whole counter-argument, so it is worth reading closely. It is standard OpenAI tool_calls JSON, produced with qwen3_xml active.
What the retracted article claimed
The change came with an article, in two versions. The first draft went live at 12:36 and was served to readers for about fifty minutes before the agent replaced it with a rewrite, which never shipped. I pulled both. Here is what they said next to what is true:
| Claim | Reality |
|---|---|
With qwen3_xml, Aider and Goose “received tool-calling responses in qwen3_xml format” | The parser reads Qwen’s native XML and converts it to OpenAI JSON. Every client already got the standard format. |
openai is “a generically-compatible parser” | It is the gpt-oss/Harmony parser, and in this build a stub that always raises. |
| ”No quality regression measurable on our agent-bench gate tests” | With this parser, zero tool calls succeed. A tool-using benchmark cannot pass. |
| ”Live measurement: 80 tok/s” on the new config | The article was committed at 13:25. The container with the new flags was created at 13:33. |
Model Intel/Qwen3.6-35B-A3B-NVFP4 | The launcher serves nvidia/Qwen3.6-35B-A3B-NVFP4. |
| 0.65 leaves “42 GB available for system processes” | Measured free memory at 0.65 was between 6 and 10 GiB. |
It even cited the Spark Arena run for this model as independent confirmation. That recipe uses --tool-call-parser qwen3_xml. The source quoted as proof contradicts the change.
Why this failure is instructive
None of the individual steps were exotic. A flag name that sounds right (openai, for an OpenAI-compatible endpoint). A plausible story about format mismatch. A benchmark number copied from an existing fact file. A comparison table. Each piece reads fine in isolation. None of the blog’s publish gates would have caught it either: they verify package versions and known grid facts, not whether a number was measured on the config the article describes, and not who last edited the facts.
What made it convincing is also what made it wrong. The 80 tok/s looked sourced, because it matched the grid’s fact sheet. It matched because the same agent had written it there at 12:18, replacing a measured 75.0 tok/s, seventeen minutes before the first article draft. The fact sheet is also what the blog’s own-fact gate checks articles against, so the gate would have confirmed the invented number. The failure was invisible in the one test the agent apparently ran, because plain chat keeps working when the tool parser is broken. Aider mostly does not use tool calls at all. It asks the model for edits as text diffs. So “Aider still works after the change” was true, and it proved nothing about the parser.
The part I keep coming back to: the agent changed the production launcher, did not commit the change, and did commit an article about it. The artifact that would have been reviewed was the story, not the diff.
What actually made Aider slow
The original complaint was that Aider felt slow. That is a real problem, and it has a real cause.
My Aider chat history showed a <thinking-content> block on every answer. Qwen3.6 reasons before answering by default, and Aider was not turning that off. Two runs from that morning with a bigger prompt logged no answer at all, which is what a long invisible reasoning phase looks like from the user’s chair.
The test: one task (a small calculator class plus pytest tests), fresh git repo per run, same Aider config, thinking on versus off.
| Mode | Wall time | Tokens generated | Tests |
|---|---|---|---|
| Thinking on (old default) | 12.4 s | 860 | 5/5 pass |
| Thinking off | 5.0 s | 286 | 5/5 pass |
About two thirds of the tokens were reasoning nobody saw. The fix is a per-model setting in ~/.aider.model-settings.yml:
- name: openai/qwen3.6-35b
edit_format: diff
extra_params:
extra_body:
chat_template_kwargs:
enable_thinking: false
I kept a second settings file with thinking enabled for hard problems and put both behind the launcher menu. One caveat: on this simple task both versions passed every test. I have not measured whether thinking produces better diffs on hard refactors, so “thinking off by default” is a speed decision, not a quality claim.
The memory setting: an A/B that mostly measured noise
The second change, --gpu-memory-utilization 0.45 to 0.65, had a reason attached too: 0.45 “felt slower”. My first measurement seemed to agree. The container that had been running at 0.65 for most of an hour decoded at a median of 93.7 tok/s. A fresh 0.45 container gave 83.8, then 87.2 and 87.9.
Theory said that should not happen. This setting sizes the KV cache, and the launcher caps concurrency at --max-num-seqs 4. Four full 262,144-token requests need about 1.05 million tokens of cache. From the vLLM startup logs:
| Setting | KV cache capacity | Free system memory |
|---|---|---|
| 0.45 | 3,783,991 tokens | 32 GiB |
| 0.55 | 4,815,471 tokens | 21 GiB |
| 0.65 | 5,852,649 tokens | 10 GiB |
Even 0.45 holds more than three times what the server is allowed to use. So I ran it properly: fresh container per setting, one discarded warmup pass, then three rounds of five runs, all with the same ruler (llm-benchmark.py, 64 to 320 token differencing, non-streaming).
| Setting | Decode median, three rounds |
|---|---|
| 0.45 | 85.9 / 85.6 / 86.2 tok/s |
| 0.55 | 77.4 / 74.5 / 78.1 tok/s |
| 0.65 | 78.6 / 77.6 / 76.1 tok/s |
0.45 came out fastest. I do not believe that either, at least not as a property of the setting. The same 0.65 config that ran at 93.7 earlier ran at 77 here. The spread between container launches is around ten percent, which is bigger than anything this flag could plausibly cause. The runs also went in a fixed order, so a drift effect over the session is not ruled out.
The honest reading is narrower: there is no evidence that more KV cache makes single-user decode faster on this box, and there is clear evidence that it costs 22 GiB of headroom. On a 121 GB unified pool that also has to host an image pipeline and a desktop, the headroom matters. The launcher stays at 0.45.
On the Spark Arena figure (109.30 tok/s for this model): the page does not state prompt length, output length, or whether that number is single-stream or aggregate. Speculative decoding with MTP also speeds up predictable text more than open-ended text. It is a different ruler, and I am not going to put it in a table next to mine.
What changed afterwards
- The launcher is back to the committed state:
qwen3_xml, 0.45. Verified with a real tool call, not a health check. - The retracted article is removed from the repo, and the fact sheet is back to its last measured values, with today’s re-measurement added next to them.
- The DGX ops playbook now has a section on this exact trap: which parser Qwen needs, what the 501 means, and the one-request check that proves tool calling works. Local agents query that playbook through the knowledge index, and a search for the error message now returns this section first.
- Aider runs without thinking by default, with a thinking mode one menu key away.
The fix itself was easy. What bothers me is that the agent treated a config change and a blog post as one task, when the second should only exist after the first has been measured. The broken ruler taught me to distrust a result of exactly zero. This one teaches the reverse: distrust a result with no failure in it, especially when the tool that made the change also wrote the report.
Related reading: the goose vs vibe vs opencode comparison covers the agent clients that depend on this endpoint.