Vozonda Is Open Source: Turn Your Sources Into Your Own Podcast
Quick take: Vozonda is the podcast app I have been building on this grid, and from today it is open source under the MIT licence. You give it articles, papers, PDFs, screenshots, YouTube videos, RSS feeds or your own notes; it writes a script with a local language model, voices it, and puts the episode into your own podcast feed. Everything can run on your own machine. The code is on GitHub, samples and a blind test against NotebookLM are on vozonda.com.
Why I Built It
I liked the audio overviews in Google’s NotebookLM from the first episode. What I did not like was everything around them: my sources go to Google, the result is a download instead of a feed, and I cannot read or change what the hosts are going to say. I wanted the same idea as a tool I own: private sources, my own podcast feed, and a script I can check before anything is recorded.
That is the whole pitch. Vozonda is not a cloud service with an open-source label on it. The open-source version is the product.
What It Does Today
You collect sources in a tray: links, PDFs, images, YouTube videos with subtitles, RSS feeds and notes, up to ten at a time. Each one can carry the episode or serve as background. A long paper that does not fit the model’s context is condensed rather than cut off, so the end of a paper is not silently lost.
Then you pick a show. There are fourteen templates, from a deep dive and a quick brief to a debate, a critique, a morning dispatch of your feeds, a feature story and a roundtable with three voices. An episode can have one to three voices or a single narrator, in ten languages, and Vozonda can translate your sources into the language you want to hear.
Before anything is voiced you can read the script and fix it. I use that step on every episode I publish. A local model writes fluent dialogue, and it still gets a detail wrong now and then; the behind-the-scenes article has two real examples.
Every episode lands in a private RSS feed with chapters and transcripts in the Podcasting 2.0 format, so any podcast app plays it. The episode page adds takeaways, highlights and shareable clips. Watchlists turn a set of RSS feeds into a daily or weekly digest that fills the feed on its own.
Local First, With Choices
The script comes from any OpenAI-compatible model server: Ollama, vLLM, LM Studio or llama.cpp. The voices run on the CPU by default with Kokoro, so no GPU is required. With an NVIDIA GPU you can switch to Qwen3-TTS, and Voxtral is available as an optional cloud engine. The voice engine is a setting, not a rewrite, and better engines will keep arriving.
Uploaded files never go to a cloud model, by design. There are no accounts and no telemetry.
For Podcasters and for Agents
Each show can stay private, go to podcast apps, or be published on Nostr as well, with an explicit confirmation before anything goes public. Listeners can zap episodes with Lightning, or stream a few sats per minute while they listen. Once you add your Lightning address, your feed carries a value-for-value split that you set freely between you, the author of the source and the app.
For automation there is an HTTP API with webhooks and an MCP server, so an agent can turn what it finds into episodes on a schedule.
What It Does Not Do Well Yet
Episodes are shorter than NotebookLM’s, seven to eleven minutes for the three test sources against eighteen to twenty-one for NotebookLM on the same sources, and the dialogue is less smooth. The local model sometimes invents a detail, which is why the review step exists and why I would not skip it. The quickstart is new; if it breaks on your machine, that is the most useful issue you can open.
Compare It Yourself
I put the features of eight tools side by side on the comparison page: NotebookLM, ElevenLabs GenFM, Jellypod, Wondercraft, Audioread, and the open-source projects Open Notebook and Podcastfy. The cloud tools are smoother and need no setup; Vozonda is the one you can own.
The blind test has two parts. In the script test you read the same part of two episodes about the same source, as plain text, and pick the one that explains it better. In the voice test you hear a minute of each. Default settings on both sides, every vote published.
How It Was Built, and Where I Bent My Own Rule
Vozonda was built by the whole agent fleet on this grid, the same fleet I wrote about in The Fleet Kept Finding Bugs. So I Froze It.. Two local Qwen models on the DGX Spark worked next to about thirty free-tier cloud models that picked up tasks from the Gitea issue queue, wrote code and reviewed each other.
I also have to be honest about something that does not fit this blog’s usual line. Two of the most important helpers ran in the cloud: Claude, through Claude Code, and Antigravity (agy), Google’s coding agent. They coordinated the fleet, reviewed its work and wrote code themselves; 154 commits in the archive name Claude as co-author. That goes against the sovereign principle I write about here, and I did it on purpose. I wanted to find out how far one person can get when the best tools available today work together with a local fleet, not to build the most ideologically pure version of something smaller.
What matters for you: the cloud helped build the tool, the tool does not need the cloud. No source, script or episode of yours ever touches those services. Vozonda runs on your own hardware with a local model, and that is the version you get.
It took a lot of time and effort. From the first commit on 21 August 2026 to the release took seven weeks: 1,220 commits and 420 tracked issues. If Vozonda is worth something to you, the support page has a Lightning address and other ways to give back, none of which need an account.
I am proud of what came out of it. Nostr and open source gave me most of what this grid runs on, and Vozonda is my way of giving something back.
Try It
git clone https://github.com/Vozonda/vozonda.git
cd vozonda
cp .env.example .env
docker compose up -d
The quickstart explains the model server and the hardware options. Releases and their notes are in the changelog, the plans in the roadmap.
What a First Run Looks Like
On launch day, 8 October 2026, I ran the quickstart once more from a fresh clone, exactly as written, because a quickstart that only works on the developer’s machine is worth nothing. This is what happened:
docker compose up -dbuilt both images in about 10 minutes, most of it downloading base images.- The health check answered at once, and
/doctorshowed the script model reachable and the Kokoro and Piper voices ready. - The first episode, from the Wikipedia article on Nostr, took 3 minutes end to end: 24 seconds for the script, 139 seconds for the voices on the CPU, 6 seconds to master. The result was a 7.9-minute MP3.
One caveat on those numbers: the script came from the 35-billion-parameter model on my DGX Spark, which is fast. With a small model on a laptop, the script step takes longer, while the voices stay on the CPU either way.
The run also caught a real bug. Wikipedia answered 403 to the Docker image until the user agent carried a contact address, which is what Wikimedia’s robot policy asks for. My own machine never saw it, because its newer TLS stack was not flagged. The fix went in before the repository went public, and that is the reason I test from a clean clone instead of trusting my own setup.
A hosted version is an option for later, paid per episode with Lightning and without accounts, if enough people want it. The waitlist is one anonymous click.
Build It With Me
Vozonda is MIT licensed and open for contributions. Bug reports, a new show format, a better voice engine, a translation, a fix to the quickstart on hardware I do not own: all of it helps. The repository on GitHub has good first issues to start with, and Discussions for ideas and questions.
Vozonda is on Nostr as _@vozonda.com. Boosts over Lightning go to vozonda@rizful.com, or pick an amount on the support page.