Free, Unlimited AI Voice Cloning With Fish Audio + Claude Code
A step-by-step build of a voice-generation web app using Fish Audio's free S2.1 Pro model, wired into Claude through MCP, with emotion control and voice cloning.
September 7thElevenLabs introduces Eleven v4, a text to speech model that reads a script for context like a voice actor, and Eleven v4 Turbo, the same model tuned for real time voice agents.
Eleven v4 reads a script for context the way a voice actor does, while Eleven v4 Turbo trades a little of that nuance for about 100ms latency so the same expressive model can run inside live voice agents.
ElevenLabs releases Eleven v4, a text to speech model built on a new architecture that reads a script's context the way a voice actor would, tracking who's speaking, what just happened, and how a line should land, so multi-speaker dialogue gets real overlap and timing even without audio tags. Eleven v4 Turbo is the same research tuned for about 100ms latency so live voice agents can use it in real time. Voice cloning holds up better across long-form regenerations, context stitching keeps a regenerated line consistent with the chapter around it, and language coverage improves most in Japanese, Brazilian Portuguese, Mandarin, and Cantonese. It's live now, free to try for 11 days.
Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.
Create a free account →
ElevenLabs states the pitch directly: v4 is the most expressive TTS model yet, built on a new architecture that reads tone, pacing, emotion, character, and context, with a low-latency Turbo sibling alongside it.

Both models share the same research; the only difference is latency. v4 targets produced content (narration, character work, dubbing) across ElevenCreative, ElevenAgents, and the API. Turbo targets real-time voice agents at about 100ms median latency.

v4 reads a script the way a voice actor would, tracking who's speaking and what just happened. A scripted two-person exchange demonstrates overlap and timing without relying on audio tags, then the video shows how stacked audio tags direct delivery on top of that.

Instant and Professional Voice Clones work identically across v4 and Turbo, but v4 holds speaker identity far more reliably across regenerations and long-form work. Existing PVCs need to be manually retrained on v4. Context stitching then keeps a regenerated line consistent with the rest of its audiobook chapter.

v4 covers 90+ languages with the biggest gains in Japanese, Brazilian Portuguese, Mandarin, and Cantonese, plus better IPA support. The video closes by pointing agent builders back to Turbo for its bidirectional streaming and real-time latency, then directs viewers to sign up and try v4 free for 11 days.
Eleven v4's gains come from reading a script for context before generating it, not from adding more manual control, and that shift is worth understanding even outside ElevenLabs.
“Don't tell anyone, but we're breaking into the world's biggest candy store.”
“11v4 was designed to read a script the way a voice actor reads one.”
“If you hadn't noticed yet, this entire video is my 11v4 voice, made with an avatar entirely inside 11creative.”
See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.
ElevenLabs opens with a direct claim: this is its most expressive text to speech model yet, and the on-camera host saying it is an AI avatar running on the new voice itself.
Same underlying research split into two latency tiers so the model fits either a production pipeline or a live conversation.
“Click the first link in the description, create a free account and pick 11v4. For the next 11 days you can use 11v4 for free.”
Direct, time-boxed free-trial CTA delivered by the same AI avatar host, paired with an on-screen 'Eleven days' title card and a description link.
00:01
00:04
00:08
00:11
00:14
00:19
00:22
00:25
00:28
00:31
00:35
00:39
00:42
00:46
00:49
00:51
00:56
00:59
01:01
01:05
01:08
01:09
01:14
01:17
01:21
01:24
01:26
01:31
01:34
01:38
01:41
01:44
01:48
01:53
01:54
01:59
02:02
02:05
02:08
02:11
02:14
02:17
02:22
02:24
02:28
02:32
02:34
02:37
02:41
02:45
02:47
02:51
02:55
02:59
03:02
03:04
03:08
03:11
03:13
03:18
03:21
03:24
03:26
03:32
03:35
03:36
03:41
03:45
03:48
03:51
03:55
03:59
04:00
04:05
04:08
04:11
04:15
04:17
04:22
04:24Add Modern Creator as a preferred source and Google shows you more of our breakdowns in Search, Top Stories, and AI Overviews. It only changes what you see, and you can undo it in your Google settings anytime.
A step-by-step build of a voice-generation web app using Fish Audio's free S2.1 Pro model, wired into Claude through MCP, with emotion control and voice cloning.
September 7thBob Doyle tests whether Fish Audio's typed emotion tags actually direct a voice performance, then proves it across five full AI-generated sketch scenes.
August 26thA full walkthrough of VoiceBox, the free open-source app that clones voices and generates AI speech entirely on your own computer.
June 23rdA full screen-share build of a self-hosted n8n automation that scrapes viral clips, generates an AI avatar video with a cloned voice, and publishes it to eight platforms — narrated step by step by the creator who built it.
March 3rdA self-declared tester of 200+ AI tools narrows the field to five worth paying for — and one bonus pick that undercuts the priciest one.
March 24thOpenAI turns ChatGPT from an answer engine into an agentic coworker — a new Work mode, a desktop app that operates your files and apps, and shareable AI-built websites, all riding on the GPT-5.6 model family.
July 9th