Fish Audio's Emotion Controls Are Better Than I Expected
Bob Doyle tests whether Fish Audio's typed emotion tags actually direct a voice performance, then proves it across five full AI-generated sketch scenes.
August 26thA step-by-step build of a voice-generation web app using Fish Audio's free S2.1 Pro model, wired into Claude through MCP, with emotion control and voice cloning.
Fish Audio's free, unlimited S2.1 Pro voice model, connected to Claude through an MCP connector, replaces paid tools like ElevenLabs for emotion-controlled text-to-speech and voice cloning.
Fish Audio offers a free, unlimited text-to-speech model (S2.1 Pro) that rivals ElevenLabs, including voice cloning and bracket-tag emotional control like [angry] or [confident]. The video connects Fish Audio to Claude Desktop through an MCP connector, then uses a single Claude Code prompt to scaffold a simple web app for testing the API. After generating an API key and dropping it into the project's .env file, the builder tests emotion-tagged generations, fixes a bug where emotion tags weren't reaching the prompt, clones a voice from a 30-second recording, and browses Fish Audio's public voice library to find and reuse community-submitted voices by name.
Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.
Create a free account →
Direct comparison hook: Fish Audio is free, unlimited, and clones voices with emotional control.

Demonstrates the same phrase delivered angry, nervous/scared, confident, and in a cloned voice.

Creates a Fish Audio account as a prerequisite for the MCP connection.

Adds Fish Audio as a custom MCP connector in Claude Desktop settings, authorizes it in the browser, and restarts the app.

A single Claude Code prompt scaffolds the web app; a separate API key is generated in the Fish Audio developer dashboard.

Pastes the key into the project's .env file, saves, and tells Claude the key was added so it can verify.

Tests the generated app, finds the emotion selector wasn't injecting the tag into the prompt, and has Claude fix it.

Drops in a short voice recording and asks Claude to clone it and add it as a selectable option.

Explores built-in default voices and Fish Audio's public community voice library, searching by name.

Shows additional vocal effects like laughing, then closes with a CTA to try Fish Audio and subscribe.
Fish Audio's free S2.1 Pro model plus an MCP connector turns Claude into a voice-app builder, but the real skill is emotion-tag prompting and treating credentials, testing, and voice cloning as separate steps.
“Forget about 11 Labs. This AI voice tool is just as good, if not better, and it's free and unlimited.”
“Claude is going to go ahead and use fish audio to clone the voice and add it as an option for future generations.”
See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.
The pitch is a direct swap: skip ElevenLabs and get free, unlimited AI voice generation with emotional control by wiring Fish Audio into Claude through an MCP connector.
Placing one or more bracketed emotion keywords at the start of a text-to-speech prompt changes the model's vocal delivery of the same underlying sentence.
“You can go ahead and try out fish audio for free with the link in the description. Be sure to subscribe for more.”
One clean CTA at the very end, paired with a subscribe ask. No hard sell mid-video beyond a joking aside.
00:01
00:10
00:15
00:23
00:30
00:36
00:43
00:50
00:56
01:03
01:11
01:16
01:23
01:30
01:37
01:43
01:50
01:57
02:03
02:10
02:17
02:23
02:29
02:36
02:40
02:50
02:57
03:03
03:08
03:17
03:23
03:31
03:37
03:44
03:50
03:57
04:04
04:10
04:17
04:25
04:30
04:37
04:44
04:50
04:57
05:04
05:10
05:17
05:24
05:31
05:39
05:44
05:51
05:57
06:04
06:11
06:17
06:24
06:31
06:37
06:44
06:51
06:57
07:05
07:11
07:18
07:24
07:31
07:38
07:44
07:51
08:01
08:04
08:11
08:18
08:24
08:33
08:38
08:46
08:51Add Modern Creator as a preferred source and Google shows you more of our breakdowns in Search, Top Stories, and AI Overviews. It only changes what you see, and you can undo it in your Google settings anytime.
Bob Doyle tests whether Fish Audio's typed emotion tags actually direct a voice performance, then proves it across five full AI-generated sketch scenes.
August 26thBen Senescu built OpenSEO to kill his own Semrush bill, wired it straight into Claude, and now watches open-source contributors fork it into CRMs and content tools he never planned to build.
August 24thA screen-share walkthrough of wiring a brand-aware Claude Code project to Higgsfield's AI models so Claude does the prompting, generates a full asset suite for a fictional energy drink brand, and reports back what each piece cost.
August 21stA creator wires Claude Code into DaVinci Resolve through an open-source MCP server, then stress-tests whether it can actually follow marker-based, relative instructions the way a real junior editor would.
July 17thA full walkthrough of VoiceBox, the free open-source app that clones voices and generates AI speech entirely on your own computer.
June 23rdA cost breakdown and two Claude Code skills for routing every coding task to the right model instead of defaulting to the most expensive one.
September 7th