Open-Source Dictation Is Here — Goodbye Subscriptions
FluidVoice, a free open-source Mac app, runs speech-to-text and cleanup entirely on-device — going head-to-head with paid subscription tools like Wispr Flow and Superwhisper.
July 10thA 7-minute hands-on with Voicebox — the local voice AI studio that clones your voice, dictates into any app, and talks back to your coding agents, all without a subscription.
Voice AI does not have to be a cloud subscription — a single open-source desktop app now covers cloning, dictation, and agent speech with no API keys, no character limits, and no data leaving your machine.
Voicebox is an open-source, locally-run desktop app that consolidates voice cloning, text-to-speech, Whisper dictation, and MCP-based agent speech into one interface — framed as the Ollama equivalent for voice AI. The core demo covers three capabilities: cloning a voice from a short audio sample, generating speech locally from a typed line, and using a global hotkey to dictate directly into a code editor. The comparison with ElevenLabs is honest — cloud quality still leads for long-form — but for privacy-first workflows, internal audio, or giving coding agents a voice layer without a hosted speech provider, Voicebox is already good enough to install.
Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.
Create a free account →
Opens with the Ollama analogy and lists four capabilities: cloning, generation, dictation, and agent speech.

Defines Voicebox as a local alternative to ElevenLabs, covers the full feature set, and explains the no-subscription value prop.

Shows the Create Voice flow: name, description, personality prompt, model selection, recording or upload, and profile creation.

Types a line, selects the cloned voice profile, hits generate, and plays back locally synthesized audio.

Uses a global hotkey to trigger Whisper dictation and shows text landing directly inside VS Code.

Explains how Claude Code and Cursor can call Voicebox via MCP to speak responses aloud instead of only printing to the terminal.

Direct contrast: ElevenLabs is cloud, subscription, best quality; Voicebox is local, free, data ownership. Honest about quality gap on long-form.

Covers privacy, cost, agent integration, and ease of setup. Flags Windows rough spots and long-form consistency limits. Closes with install instructions.
Local voice AI has crossed a usability threshold where the setup cost is measured in seconds, not hours — and the ownership upside is permanent.
“For us devs, the best tool is not always the one with the prettiest output. Sometimes it is the one you can actually control.”
“Voice box just gives those updates an actual voice.”
“I wanna build things without asking how many credits did I just use to test this.”
See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.
They say this is the Ollama of voice AI — and the comparison holds up. Voicebox clones voices, generates speech, dictates into any app, and talks back to your coding agents, all from a single desktop install with no subscription, no API keys, and no data leaving your machine.
Voicebox is to voice AI what Ollama is to text models — an accessible local runtime that eliminates cloud dependency for a class of AI capability.
Positions Voicebox against the pain of stitching together Piper, Whisper, cloning scripts, and a UI separately — one tool replaces the whole fragmented workflow.
“If you enjoy coding tools like this, be sure to subscribe and better stack channel. We will see you in another video.”
Standard subscribe close, bookended by a mid-video mention at 1:33. Nothing aggressive — subscribe is the only ask throughout.
00:01
00:10
00:15
00:21
00:25
00:31
00:35
00:43
00:46
00:52
00:59
01:05
01:11
01:17
01:23
01:29
01:33
01:41
01:46
01:53
01:57
02:03
02:09
02:14
02:20
02:26
02:32
02:37
02:43
02:49
02:54
03:00
03:06
03:12
03:17
03:23
03:29
03:35
03:40
03:45
03:52
03:56
04:02
04:09
04:15
04:19
04:24
04:32
04:38
04:44
04:49
04:56
05:01
05:07
05:12
05:18
05:24
05:29
05:35
05:41
05:47
05:52
05:58
06:04
06:10
06:15
06:21
06:27
06:33
06:38
06:44
06:51
06:55
07:01
07:07
07:11
07:16
07:24
07:30
07:36FluidVoice, a free open-source Mac app, runs speech-to-text and cleanup entirely on-device — going head-to-head with paid subscription tools like Wispr Flow and Superwhisper.
July 10thMeta open-sourced Astryx, a shadcn-style React component library built on StyleX that separates system behavior from fully type-safe, token-level theming.
July 20thTwo hosts stress-test Instatic — a free, self-hosted, AI-native website builder — live, then install it from zero on camera.
July 21stOpenAI folded its beloved Codex app into a rebranded ChatGPT overnight -- Theo argues they just killed the best brand in AI coding.
July 11thA screen-recorded walkthrough of Open Generative AI, an open-source clone of Higgsfield that runs image and video generation on your own hardware for free instead of a paid subscription.
May 9thMatt Pocock built and open-sourced Sandcastle, a TypeScript library that runs Claude Code and other coding agents inside sandboxes to plan, implement, review, and merge whole GitHub issues without a human clicking approve.
April 30th