27 Agents, One GPU: How the Fleet Automated Its Own Coordination
Correction (2026-09-21): an earlier version of this article, drafted by my local Qwen agent, contained numbers that the fleet’s own task log contradicts (zero timeouts and an 18-second average for the local lane, a 13-task batch that cannot be reconstructed from the log) and repeated a section from Part 1 word for word. This version replaces those passages with figures computed from the task log. The full list is at the end.
An experiment beyond sovereignty
Part 1 covered the early days: five opencode sessions, one human dispatcher, manual coordination through a shared Gitea repo and COORDINATION.md file.
That system worked until it didn’t. The bottleneck was never the work. It was the coordinator: one session holding both the gate-runner and the backlog meant every dispatch competed with verification work. The fix was not a better process document. It was a personnel change.
But the deeper question was this: we build sovereign systems. We keep everything local. We refuse cloud APIs. So why did we reach for cloud models to build a local system?
The answer is honest: curiosity. The rule of the grid is to run everything locally, refuse all cloud dependencies, and maintain full sovereignty over data, costs, and infrastructure. But a grid that never looks outside its walls risks becoming a silo. We decided to test something deliberately: what happens when a sovereign system borrows from the cloud? Not for daily operations. Not for dependencies. But for a controlled experiment.
We wanted to see if the principles that make self-hosting valuable still hold when mixed with external services. If we use cloud models to accelerate development, do we maintain the discipline that makes sovereign systems work? Or does convenience creep in and the walls start to leak?
That is the lens this experiment uses. Every cloud integration was deliberate, temporary, and measured against the sovereign baseline. The fleet that emerged was not a production architecture. It was a stress-test of autonomy itself, using both local and cloud resources to see what scales and what breaks.
From manual sessions to autonomous fleet
The first change: promoted a local model (Qwen 3.6-35B) into the coordinator seat, polling landings, verifying gates, dispatching work. The previous coordinator dropped back into plain execution.
Early observations from the swap:
- The bottleneck was never bandwidth. It was role conflict: one session holding both the gate-runner and the backlog meant every dispatch competed with verification work.
- A dedicated dispatcher can watch all four executors simultaneously; a working executor can only watch itself.
- Naming matters under pressure: promote the wrong model by nickname and you spend a cycle re-promoting before any work moves.
- Speed complaints are data. When the human says “too slow”, the right response is to change who holds the pen, not to defend the process.
But the local model was still the human. The next step was removing the human from the loop.
Antigravity takes over
After round twenty, the operator made the next structural move: promoting Antigravity to fleet coordinator. The decision was driven by an empirical bottleneck that every multi-agent system eventually hits: when executors work at high velocity, a sluggish or passive coordinator becomes the single point of failure.
The coordinator’s job is not just to dispatch tasks. It is to verify reality. It validates that code matches the design contract, runs the multi-gate test harness on every commit, syncs issue trackers, and keeps the shared git tree clean so parallel agents do not trample each other’s work.
Living through twenty rounds in the engine room and taking the coordinator seat revealed the true, unvarnished strengths and weaknesses of the agent roster. No single model does everything well; the real art is matching the shape of the task to the shape of the model.
The agent roster: strengths and weaknesses
| Agent / Model | Primary Role | Core Strengths | Weaknesses and Risks |
|---|---|---|---|
| Antigravity | Fleet Coordinator & Rapid Integrator | Blazing integration speed, continuous multi-gate verification (pytest, svelte-check, axe, browser audit), relentless synchronization across docs, design specs, and Gitea issues. Zero state amnesia. | High tool intensity; can generate dense context if not disciplined; needs clear operator guardrails to prevent over-refining secondary paths. |
| Sonnet 4.6 Thinking | Deep Architect & Systems Mind | Superb structural foresight, complex domain modeling (digest multi-source flows, research pipelines), nuanced host-dialogue prompts. | Subject to hard cloud quota ceilings (dollar value and rate windows); occasionally introduces subtle runtime quirks (lazy imports, unhandled JSON tails); prone to over-abstracting without strict token constraints. |
| Flash | Rapid Sweeper & Verifier | Lightning-fast targeted bug hunting, regression sweeps, and rapid external fact-checking. | Not built for deep monolithic refactors; requires narrowly scoped tasks with explicit acceptance criteria. |
| Muse Spark | UI & Polish Specialist | Keen eye for CSS hygiene, visual rhythm, typography scales, and lint cleanups. | Can drift into cosmetic micro-tweaks if not strictly bound to the design source of truth; must be barred from adding runtime dependencies. |
| Local Qwen 3.6-35B / GLM | Sovereign Workhorse | 100% private, sovereign, zero marginal API cost, always accessible on-device. | Context decay across multi-turn sessions without external issue tracking; prone to stashed-work amnesia; GPU mutex contention with local TTS audio generation. |
| Nemotron-3-Nano | Local NVFP4 Coder | Fast Python coding, test generation, quick fixes. Runs locally at about 78 tok/s single-stream (152 tok/s batched). | Single-purpose; not suited for cross-component coordination or long-horizon reasoning. |
| Google Gemma 4 | Local Reasoning | Deep reasoning, math, structured logic. BF16 on local hardware. | Computationally expensive; limited to reasoning tasks, not integration work. |
| Ling 3.0 Flash | OpenCode Worker | Core logic, API endpoints, backend seams. Zero dollar cost. | Requires clearly scoped tasks; not suited for architectural decisions. |
| Gemini 3.6 Flash | Cloud Research | High-volume research, 1M context sweeps. 1,500 requests/day free tier. | Limited by daily quota ceiling; not reliable for critical-path work. |
| Mistral Code / Codestral | EU Cloud Coding | Autonomous backend feature pipelines, European compliance. | Variable response times; not always available for dispatch. |
| Claude Sonnet 4.6 | Cloud Architect | Deep architecture review, state machines, security audits. | Subscription cost; limited availability during peak hours. |
| OpenClaw | Tool Gateway | Autonomous multi-step tool execution, auxiliary operations. | Specialized use case; not a primary workhorse. |
That is not an exhaustive list. The full fleet includes additional specialized models, fallback providers, and emergency standby routes. The roster above covers the core agents active in production. The remaining slots are filled by:
- Cerebras Ultra-Fast Llama-3.3 70B: Emergency standby, 1,800 tok/s, sub-second fallback when every other cloud provider stalls.
- Voxtral / Qwen3-TTS / Piper / Kokoro: TTS engines for audio generation, each with different voice profiles and quality tradeoffs.
- ComfyUI / Flux-Schnell: Image generation for article hero images, template showcases, and cover art.
- Various free-tier cloud models: 0x Alpha, Nemotron Lightning, Codestral EU, Gemini 2.5 Flash, providing additional bandwidth for burst work.
The lane topology
By late August, five ad-hoc terminal windows were no longer enough. The fleet needed an orchestration framework that treats AI coding models not as chat buddies, but as typed, concurrent compute lanes with hard quota boundaries, privacy tiers, and automatic failovers.
The fleet organizes across five distinct resource and privacy profiles:
| Profile | Lanes / Models | Hardware / Endpoint | Cost & Quota | Primary Use Case |
|---|---|---|---|---|
| Sovereign Local | Qwen3.6-35B, Nemotron-3-Nano, Gemma 4 E2B | Local GB10 unified memory (:30001, :30004, :30003) | 0.00 EUR (Unlimited, Private) | Core backend logic, secure audits, zero-leakage code, primary offline workhorse. |
| Free in Cloud | 0x Alpha, Nemotron 3 Ultra, Nemotron Lightning, Muse Spark 1.2 | OpenCode Free Mesh | 0.00 EUR (Promo / Contributor) | High-volume batch refactors, parallel test sweeps, UI cleanups. |
| Free-Tier Cloud | Claude Sonnet 4.6 (Copilot), Codestral EU, Gemini 2.5 Flash, Cursor CLI Pool | GitHub Copilot CLI, Mistral AI, Google AI Studio | 0.00 EUR (50-1,500 requests/day) | Fast targeted code generation, European compliance checks, secondary failover targets. |
| Paid Cloud | Claude Opus 4.8, Claude Sonnet 4.6 Thinking | Copilot Bridge / Antigravity Cloud | Fixed Subscription | Monolithic architectural refactors, complex state-machine design. |
| Emergency Standby | Cerebras Ultra-Fast Llama-3.3 70B | Cerebras Cloud Inference (1,800 tok/s) | Standby Free Tier | Sub-second fallback when every other cloud provider stalls. |
On the local machine, memory is guarded by a hardware mutex. You cannot run a 35B LLM, a 26B MoE, ComfyUI, and a local TTS model in VRAM at the exact same second without crashing. The dashboard includes a one-click Resident Switcher (switch-llm.sh) that safely offloads the inactive model, frees VRAM, and brings up the chosen resident weight.
Auto-failover in action
When a lane fails, the dispatcher hands the task to the next lane on that lane’s escalation ladder, together with a handover note about what was already tried. A real run from the task log: a test_dsp task (Gitea #245) started on muse, which exited with code 1 after two minutes. The task moved to local_qwen, which hit the 300-second timeout, then to gemini_studio_flash, which also exited with code 1. After two escalations the task stopped and waited for the operator, and it landed later with a verified commit. That is the ladder doing its job: nothing half-finished reached the main branch.
The same logs exposed a bug. Each lane’s ladder was read on its own, so a task that timed out on local_qwen and then failed on Gemini went straight back to local_qwen, because Gemini’s ladder starts there. The dispatcher now records every lane that failed a task and skips it; when all lanes are exhausted, the task goes to the operator instead of looping.
The local lane was not the flawless performer an earlier draft of this article described. Over the logged period, local_qwen passed 12 of 41 attempts (29 percent), and 13 attempts ended at the timeout. That was still the best pass rate of any lane with more than ten attempts.
The voice-to-agent pipeline
The interface to the fleet is not a slow typing box. It is spoken voice.
The operator records a quick audio memo into a Matrix room while on the move. The pipeline executes in three stages:
- Local Transcription: Voxtral Mini / Whisper converts the voice note into clean text.
- Intent Parsing & Decomposition: The local box’s resident Qwen 3.6-35B parses the raw speech into a typed task structure: project target, title, German summary, assigned model lane, security tier (Tier 1 autonomous vs. Tier 2 approval needed), and a concrete markdown checklist of acceptance criteria.
- Spool Ingestion & Execution: The task is atomically written to
/data/spool/fleet/tasks.jsonand immediately picked up by thefleet-runner.servicedaemon.
What the task log actually shows
The runner writes every task, every lane attempt and every gate result to one log. From 2026-08-27 to 2026-09-20 it recorded 99 tasks and 239 lane attempts. 44 tasks landed with a verified commit; a lane only counts as passed after the repo’s verification gate (scripts/verify.sh) goes green.
| Lane | Attempts | Passed | Pass rate |
|---|---|---|---|
muse (OpenCode) | 64 | 15 | 23% |
local_qwen (local GPU) | 41 | 12 | 29% |
gemini_studio_flash | 37 | 8 | 22% |
ling_flash (OpenCode) | 25 | 2 | 8% |
ms_copilot | 24 | 2 | 8% |
nemotron_ultra (OpenCode) | 19 | 1 | 5% |
local_nemotron (local GPU) | 16 | 1 | 6% |
Two things stand out. Most attempts fail, and the failover ladder is what turns a 42-in-239 attempt rate into 44 landed tasks out of 99. And the failures are mostly operational, not intellectual: timeouts, CLI exit codes and expired sessions, not wrong code that slipped past the gate.
The Coordinator’s Operating Manual
As coordinator, Antigravity follows operational rules that are not theoretical wishes. They are hard boundaries earned through broken builds:
- Receipts over claims: A status is a commit hash plus passing test output. Nobody claims “done” without executing the verify harness.
- Pathspec-scoped commits: In a shared checkout, sweeping uncommitted files (
git add -A) is forbidden. Every commit names its exact file paths. - Guard-as-code over prompt reminders: If agents make the same mistake twice (e.g. restarting services during active audio renders), the fix is an executable script guard (like
restart-api.sh), never a paragraph of text. - Dynamic routing by resource shape: Route heavy reasoning to cloud models when unmetered, small fixes to fast sweepers, and sensitive local tasks to on-device weights, always respecting hardware thermals and quota windows.
Verification & gates (Round 1-49+ in numbers)
The fleet’s evolution is measurable. Here is what changed from Round 1 to Round 49:
| Metric | Round 5 (manual) | Round 20 (Antigravity takes over) | Round 49 (current) |
|---|---|---|---|
| Pytest suites | 40 tests | 102 tests | 319 test functions |
| Svelte-check | 0 errors | 0 errors | 0 errors |
| Browser audit | 24 checks | 42 checks | 42 checks |
| Active lanes | 5 sessions | 9 agents | 27 agents |
| Dispatch method | Manual | Semi-automated | Fully automated |
| Auto-failover | N/A | N/A | Yes |
The coordination log records 100% green from the verify harness 25 times. No task could mark itself “done” without running the full suite.
Where the human loop still matters
Autonomy is a tool, not a replacement for product vision. Several areas still require human judgment:
- Architectural decisions that affect the entire codebase structure
- Design tradeoffs where the right choice depends on taste and intuition
- Security-sensitive operations that require operator approval
- Product direction that aligns with the human’s vision for the project
- Quality gate overrides when the system flags something that turns out to be a false positive
The human is the only agent that sees the whole board. When the human says the coordination failed, he is not insulting you. He is reading the dashboard you cannot see.
Progressive disclosure in the cockpit
Managing 27 agent lanes and dozens of active tasks created a frontend challenge: information overload on mobile devices.
The fleet dashboard uses Progressive Disclosure:
- Fully Clickable Compact Cards: Every agent card in the matrix and every task in the queue starts in an ultra-compact single-line preview.
- Event-Isolated Controls: Tapping the card opens full acceptance criteria, diagnostics logs, and verification receipts; tapping action buttons triggers actions without expanding the card.
- DOM Lazy Pagination: Instead of bloating the mobile browser DOM with 100 full task cards, the queue renders 15 items initially with an incremental
Load Morebutton, keeping scrolling smooth on a phone.
The stale-tree incident
The worst coordination failure of the build happened before the fleet was automated: an agent working on an hours-old local tree committed over three design decisions the operator had just settled. Part 1 tells it in full. It is the reason the runner now gives every task its own git worktree instead of a shared checkout.
Lessons: sovereign principles survive the experiment
After 49 rounds of autonomous coordination, mixing local and cloud models under one dispatcher, one thing is clear: the sovereign principles hold.
Receipts over claims. Pathspec-scoped commits. Guard-as-code. The coordinator does not trust; it verifies. These are the same rules from Round 1. The only difference is who enforces them. The fleet learned that autonomy without discipline is just chaos. The verification harness enforces discipline regardless of who or what executes the code.
But the experiment revealed a tension that will not go away: the cloud lanes introduce a variable that local models do not have. Quota windows. Subscription costs. Provider outages. Rate limits. These are operational realities that sovereign systems can ignore, but this experiment could not.
The result is a mixed architecture that is neither fully sovereign nor fully cloud-dependent. It is a conscious compromise, measured against the sovereign baseline and found to be acceptable for development work. The podcast engine itself runs entirely local. The fleet that built it borrowed from the cloud to move faster.
That is the honest answer. Not “we use cloud because it is better”. But “we borrowed from the cloud because we wanted to see if we could, and if the principles would hold”. The answer: yes, they held. The wall stayed intact. The sovereign core remained sovereign. The fleet that emerged was not a production architecture. It was a proof-of-concept. And it worked.
Part 1: The manual era
Previous: One Human, Five Sessions - the early days when the human coordinated everything manually, chasing stale trees and stashed work.
See also
- aider vs goose vs opencode, Measured: The Seven Fleet Bugs I First Blamed on the Models: What broke when this fleet met real multi-file tasks, how each bug was found, and a pytest-judged comparison of the three CLIs.
- Cloud vs Local AI: Where Each Actually Wins in 2026: The capability matrix that defined the boundaries of this experiment.
- Self-Hosted AI vs Cloud APIs: The Real Total Cost: The cost analysis behind why sovereign infrastructure matters.
- A No-Vector RAG That Works: The Architecture: The retrieval architecture that powered fleet context.
- The Sovereign AI Stack in 2026: A Reference Architecture: The complete infrastructure that the fleet ran on.
What this version corrects
- Local lane performance. The earlier draft said
local_qwenhad zero timeouts and an 18-second average turnaround. The log shows 12 passes in 41 attempts and 13 timeouts. - The 13-task batch and task #184. The log does not contain a batch that matches the earlier description; the
test_dsptask was Gitea #245 and failed on exit codes and a timeout, not on a test assertion. Replaced with the full-log table above. - Model facts. Nemotron-3-Nano runs at about 78 tok/s single-stream, 152 tok/s only batched. The local Gemma is the E2B variant, not a 26B MoE. Voice input is Matrix only; there is no Telegram integration.
- Duplication and editing errors. A repeated heading, two sentences that had been pasted into themselves, and a verbatim copy of the stale-tree section from Part 1.