aider vs goose vs opencode, Measured: The Seven Fleet Bugs I First Blamed on the Models
For weeks, every agent in my self-hosted coding fleet received the first 200 characters of each task and nothing else, and I read the resulting failures as proof that the models were weak.
Quick take: Twenty-two tasks on one repo, 126 lane attempts across local and cloud models, and not one landed. The cause was not model quality but seven bugs in the orchestration layer: truncated task text, worktrees that imported the main checkout’s code, aider pulling every mentioned file into a 262k context until it sent 430k tokens and got nothing back, generated files counted as scope violations, a handover that deleted the previous agent’s work, a LIFO queue, and a free-tier quota that stalled silently. After fixing them I measured aider, goose and opencode on the same models with pytest as the judge. On a local Qwen, aider solved all three tasks, goose two and opencode none, because the model wrote its tool calls as plain text. On strong free NVIDIA NIM models all three CLIs solved everything, and aider was the fastest.
The fleet is the setup from 27 agents, one GPU: a runner that gives each task its own git worktree, a sandboxed agent CLI, a gate (verify.sh plus a task test), and an escalation ladder that hands a failed task to the next lane. Claude Code acts as the coordinator that writes the task specs; the lanes are free or local models. On the day this post starts, the queue for one product repo held twenty-two tasks. Over 126 lane attempts, the fleet landed none of them; the only one that shipped, the coordinator wrote by hand.
The symptom: every lane fails, so it must be the models
The pattern looked like a capability problem. Six of eight tasks hit the 15-minute timeout. Others failed with “Task gate target missing: the test the task asked for was not created” or “Scope violation: file outside allowed_paths”. A task that asked for a registry refactor edited main.py, a file its spec explicitly forbade. It happened on a local Qwen, on Gemini, on Mistral, on NVIDIA-hosted Kimi and Nemotron.
When every model fails the same way, the constant is the harness. That turned out to be true seven times over.
Bug 1: the agent saw 200 characters of a 4,000-character task
The dispatcher stored a short summary for the dashboard, summary_de = raw_text[:200], and the prompt builder used that summary as the task description. My specs were 1,400 to 4,000 characters: file lists, the exact test path to create, which files not to touch. The agent got the first sentence and a list of acceptance criteria.
The live prompt of a running agent showed it: the documentation task stopped in the middle of an enumeration, followed directly by “Acceptance Criteria”. Every “test file not created” and most scope violations trace back to this. The agent was never told the path or the boundary.
Fix: the prompt now carries the full task text (capped at 12,000 characters), the allowed paths as an explicit hard limit, and the gate command including the test file that must exist. Prompts went from roughly 200 characters of task to 10,000 to 12,000 characters.
Bug 2: every worktree tested the main checkout’s code
Each task runs in its own git worktree, and the Python virtualenv is symlinked in to save disk and time. That virtualenv contains an editable install of the project: a .pth file whose single line points at apps/api/src of the main checkout. So inside the worktree, import hearsay_api resolved to the main checkout, not the worktree.
The agent wrote a new function, its new test imported the old code, the test failed, the agent “fixed” it again, until the timeout. Worse, all worktrees and the main checkout shared one SQLite test database, which made unrelated tests flaky whenever two runs overlapped. Running a failing test file in isolation passed 16 of 16; the same files under parallel load failed.
Fix: the runner reads the .pth files of the linked virtualenv, maps any path inside the main checkout to the same path inside the worktree, and prepends it to PYTHONPATH for both the agent and the gates. Verified on a live agent: hearsay_api now imports from the worktree and writes to the worktree’s own database.
The same bug had a second half that surfaced a day later. The repo’s committed verify.sh ran uv run --extra tts pytest without --no-sync. Inside a worktree, uv “helpfully” reinstalled the editable package, now pointing at the worktree, into the shared virtualenv. Every gate run silently rewired the main checkout’s imports to a half-finished branch; the main checkout’s own test run failed with an ImportError from a path under the fleet’s worktree directory. The fix was already sitting in the main checkout as an uncommitted edit, so every worktree kept getting the broken version. It is committed now, and the runner sets UV_NO_SYNC=1 for agents and gates as a second line of defence.
Bug 3: aider sent 430k tokens to a 262k model and exited 0
This one hid behind a clean exit code. Replaying a real task prompt through aider against the local Qwen (262,144 tokens of context) printed:
Tokens: 430k sent, 0 received.
Model openai/qwen3.6-35b has hit a token limit!
aider adds every file a message mentions to the chat, and --yes-always confirms each addition. The prompt included the repository contract, which names about twenty files, among them uv.lock. aider loaded them all, blew the window, received nothing, and still exited with code 0. The runner saw a clean exit and no diff and reported “Agent produced zero file changes”. This hit every aider-based lane at once.
Fix: the repository contract goes to aider as a read-only file (--read AGENTS.md) instead of message text, a fleet-wide .aiderignore hides lockfiles, media, databases and build output, and URL scraping is off (--no-detect-urls). The same prompt then sent 15k tokens and the edit landed in 25 seconds.
Bug 4: generated files counted as scope violations
Once agents saw the full task, they worked harder, and they left debris. A documentation task failed its scope check on generated sample audio under apps/web/public/media/samples/voices/; other attempts left a downloaded pandoc .deb, a tmp/probe-*.mp3, or an aider >>>>>>> REPLACE fragment behind. None of it would ever have been committed, and each one failed an otherwise finished task.
Fix: untracked files outside the allowed paths are deleted in the worktree before the scope check (never in the main checkout). Only tracked edits outside the scope still fail a task, which is what a scope violation is supposed to mean.
Bug 5: the handover threw away the previous agent’s work
When a lane failed, the runner removed the worktree and the branch, and the next lane started from a clean checkout with a 1,200-character diff excerpt as “context”. An agent that had finished most of a task and failed on one lint rule handed nothing usable to its successor.
Fix: a failed attempt with changes is committed as wip(<lane>), its worktree is kept, and the next lane resumes it with an explicit note to continue rather than restart. A red gate now also gets one repair pass on the same lane with the failure output before escalating. One follow-up bug came with it: the “zero changes” check looked for a dirty tree or a commit in the last five minutes, so a resumed branch with inherited commits failed as “no changes”. A branch ahead of main now counts as work.
Bug 6: the queue ran newest first
The task file is stored newest first and the runner iterated it as-is. A documentation task queued last overtook security fixes queued an hour earlier, and dependent tasks could start before their prerequisite. The fix is a sort by creation time; an escalated task keeps its original timestamp, so its retry still runs before later work.
Bug 7: a free-tier quota stalled silently until the timeout
The Gemini free tier ran out of its daily quota. opencode’s log showed You exceeded your current quota, but opencode kept retrying with backoff until the task timeout, so each attempt burned the full slot. The runner now recognises the daily-quota message, blocks that lane until the quota resets at midnight Pacific time, and skips it immediately. Attempts that fail for infrastructure reasons (runner restarts, a misconfigured gate, quota) are logged as [infra] and no longer count against the lane’s measured pass rate, which had been dragged down by failures that were never the model’s.
aider vs goose vs opencode on a local Qwen and free NVIDIA NIM models
With the plumbing fixed, the question I had dodged in June was answerable: which coding CLI actually does the work? I ran the same three tasks through each CLI, via the fleet’s own sandboxed helpers, in throwaway git repos, and let pytest decide instead of the agent’s own summary.
- t1 create: write
ping.txtcontainingpong. - t2 bugfix: a failing test expects
div(1, 0)to raiseValueError; fix it and change onlycalc/ops.py. - t3 feature: add
calc/stats.pywithmean()that uses the existingdiv, plus at least two tests, and touch nothing else.
The models: the local Qwen3.6 35B MoE on vLLM, and two free NVIDIA NIM endpoints, Kimi K3 and Nemotron 3 Ultra. Timeouts were 8 minutes for the local model and 20 minutes for NIM, whose free tier queues requests.
| Model | aider | goose | opencode |
|---|---|---|---|
| Qwen3.6 35B (local) | 3/3 (5 to 8 s) | 2/3 (29 to 35 s) | 0/3 |
| Nemotron 3 Ultra (NIM) | 3/3 (9 to 27 s) | 3/3 (45 to 63 s) | 3/3 (65 to 100 s) |
| Kimi K3 (NIM) | 3/3 (3 to 11 min) | 3/3 (11 to 20 min) | 3/3 (14 to 20 min) |
Four of the Kimi runs (two each for goose and opencode) were still running at the 20-minute limit; the judge ran after the kill and found the work complete, so they count as passed but their times are really “at least 20 minutes”.
Three observations.
The CLI only matters for the weak model. With Nemotron Ultra and Kimi K3 every CLI solved every task; the differences are speed. With the local Qwen, opencode solved nothing: the model wrote its tool calls as plain text (echo -n 'pong' > ping.txt printed, not executed; <Glob and leftover <think> tags in the output), so no edit ever happened. That matches what I found earlier about Qwen’s tool calling. aider does not need structured tool calls at all; it asks for search/replace blocks in plain text and applies them itself, which is exactly why it survives a model with shaky tool calling.
aider is the fastest everywhere. One request, one edit, done. goose and opencode run an agent loop, read files, run the tests, then answer, which costs minutes when every request waits in a free-tier queue.
Free NIM endpoints come and go. In the first run Nemotron Ultra answered HTTP 404 through all three CLIs; when I probed it again later that morning it answered 200. Treat free endpoints as best effort: my fleet probes them daily and only marks a lane dead on a permanent status, never on one bad hour.
Two mistakes of my own are in these numbers, corrected. My first t3 started from the repo with the deliberately failing t2 test, so “all tests pass” silently required fixing an unrelated bug; goose did that and was rewarded, aider stayed in scope and was punished. And goose creates a private .venv to install pytest, which my scope rule first counted as a violation. t3 now starts from a green repo, and tool scratch like a .venv is ignored, the same way the fleet now cleans it up. One more goose detail: it loads the operator’s own MCP extensions from ~/.config/goose, including a broken one in my case, so give a fleet its own goose config.
The sample is small, one run per cell on toy tasks, and NIM timings depend on queue load. It says which tools can do the job on which model, not which is best at scale.
Is aider still a safe bet in 2026?
The numbers favour aider, the maintenance record does not. As of 2026-09-23, aider’s last release is v0.86.0 from August 2025 and its last commit is from May 2026, with close to 1,900 open issues. goose moved to the Agentic AI Foundation and shipped v1.51.0 on 2026-09-17; opencode shipped v1.18.32 on 2026-09-21. aider still works today, but a CLI that nobody updates will eventually break on a new model API.
Today the local Qwen lane and all five NIM lanes in my fleet run aider, and the local lane falls back to goose. goose is the tool I am moving the longer, multi-step NIM work to next, where its agent loop pays off and its maintenance is not a question. Routing is measured per lane, so if aider starts failing, work moves away from it on its own.
A gate caught a crypto bug through a lint rule
One more result from the same day. A fleet task had to add BIP-340 signature verification to a Nostr login. The first attempt that got through all the plumbing failed only on ruff: F841 Local variable e is assigned to but never used. In Schnorr verification, e is the challenge the whole check depends on; an unused e means the code accepted signatures without verifying them. A style gate stopped a login that anyone could have passed. I wrote that function by hand from the reference implementation and tested it against the 19 official test vectors.
A checklist for anyone running coding agents in worktrees
- Print the exact prompt an agent receives, from the live process (
/proc/<pid>/cmdline), before you judge the model. - If worktrees share a virtualenv with an editable install, set
PYTHONPATHto the worktree’s source or every test runs against the wrong code. - Replay one real task through each CLI and read the token count. A clean exit code with zero output tokens is a context overflow, not a lazy model.
- Treat untracked build and test artifacts as cleanup, and only tracked edits outside scope as violations.
- Keep a failed attempt’s work for the next agent.
- Exclude infrastructure failures from any per-model success rate, or you will route work away from good models for bad reasons.
The earlier decision post, goose vs vibe vs opencode, picked a primary CLI without a fixed task suite. The bake-off above is that missing measurement, built on the method from agent-bench: small tasks, deterministic gates, numbers instead of impressions.