Fish Audio's Emotion Controls Are Better Than I Expected
Bob Doyle tests whether Fish Audio's typed emotion tags actually direct a voice performance, then proves it across five full AI-generated sketch scenes.
Posted
3 days ago
Duration
Format
Demo
educational
Views
2.4K
72 likes
57 · 43
Big Idea
The argument in one line.
Fish Audio's free-text emotion tags turn text-to-speech into directable voice acting, but they only prove their worth once a tagged line pairs a consistent cloned voice with real picture and sound, not when judged as an isolated audio clip.
Who This Is For
Read if. Skip if.
READ IF YOU ARE…
You make videos with AI voiceover or AI characters and want your lines to sound directed instead of flat and robotic.
You're evaluating text-to-speech platforms and want to see whether an emotion-tag system actually changes a performance or is just a marketing checkbox.
You build recurring AI characters for animation, shorts, or a series and need a way to keep one voice consistent across many emotional states.
SKIP IF…
You want a plug-and-play automated pipeline — this is a manual demo of typing tags and clicking regenerate inside one platform's UI, not a scripted workflow you can copy directly.
You're looking for an agnostic platform comparison — the video is a sponsored partnership with Fish Audio and only demonstrates that one tool.
TL;DR
The full version, fast.
Bob Doyle tests Fish Audio's emotion-tag system for AI voice, proving the same script, voice, and character can sound like three different performances just from a bracketed tag like [delighted] or [scared, whispering]. He walks the Text-to-Speech panel's free-text tags and Auto Tag AI suggestions, a wording tip that swapping formal phrasing for contractions makes lines sound more natural, then Story Studio for multi-line scripted scenes, and Instant Voice Clone for building a consistent 'stable' of emotional variants for one recurring character. His real argument: judging AI voice on audio alone is misleading, since picture and sound effects are what make a tagged line actually land, shown across five full AI-generated sketch scenes.
Free for members
Chat with this breakdown — free.
Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.
The line 'I can't believe you did that' is played three ways with only the bracketed emotion tag changed. States the real question of the video: how much control do emotion tags give over a character's performance, and does it actually change how the line lands.
00:26 – 04:19
02 · Emotion tags inside Text-to-Speech
Walks the Fish Audio TTS panel: a flat, tag-free line versus a tagged one; typing free-text tags like [delighted], [disgusted], [scared, whispering] instead of only picking from the preset list; switching to a British voice for [nervous] and [nervous, frustrated] reads; the Auto Tag button that reads a line's context and inserts tags, pauses, and emphasis automatically; a wording pro-tip swapping 'you have got' for the contraction 'you've got'; layering a [giggle] effect tag on top of an emotion tag.
04:19 – 06:34
03 · Story Studio and the cooking-show scene
Introduces Story Studio for pasting a full multi-line script, assigning voices, and adjusting timing on a timeline. Walks through a fake cooking-show scene where a saucepan's sauce becomes sentient and starts talking, showing the emotional arc line by line and how quickly a bad take can be regenerated.
06:34 – 08:55
04 · Instant Voice Clone and the stable-of-characters concept
Shows Instant Voice Clone used to build two custom character voices from scratch, a confident airline-captain voice and Bob's own cloned voice, then explains building a 'stable' of emotional variants of one character so the voice stays consistent. Closes with the argument that judging TTS on audio alone is misleading, since picture and sound effects are what sell a real performance.
08:55 – 18:05
05 · Sketch showcase and full scene playback
Quick-cut teasers of five AI-generated sketches with their live emotion tags overlaid (medieval knighthood training, an AI-cloned girlfriend, a supervillain's Airbnb'd secret lair, a Mars mission missing its oxygen tanks, plus the cooking-show sauce monster), a subscribe CTA, then all five scenes replayed back-to-back in full.
Atomic Insights
Lines worth screenshotting.
Fish Audio's emotion tags let you type natural-language direction like [delighted] or [scared, whispering] directly into the text box instead of picking from a fixed preset list.
The same script, voice, and character can produce three completely different emotional performances just by changing the bracketed tag typed in front of the line.
Fish Audio's Auto Tag button reads the context of a line and automatically suggests emotion tags, pauses, and emphasis when you don't know what to type.
Swapping the formal phrase 'you have got to' for the contraction 'you've got to' made a generated line sound noticeably more natural, independent of any emotion tag.
Combining tags like [nervous, frustrated] produces a different, more specific read than either tag alone, the same way a director layers notes for a real actor.
Story Studio lets you paste a full multi-line script, assign voices per line, and adjust timing on a timeline to build an entire audio production, not just single lines.
Instant Voice Clone lets you record exactly how you want a character to sound and turn it into a reusable voice, instead of searching a stock library for a close match.
Building a 'stable' of cloned voice variants, like a calm captain and a panicked captain, keeps a recurring character's dialect, pacing, and rhythm consistent across every scene.
Judging AI voice realism from audio alone is misleading, since real videos combine voice, sound effects, and picture to sell a performance.
Fish Audio's tag library includes non-emotion performance effects too, like laughing, chuckling, moaning, and sobbing, layered on top of the emotional read.
Regenerating a single flawed line, like a mispronounced word, takes seconds inside Story Studio, making iteration on a bad take nearly free.
Takeaway
Emotion tags only pay off when the whole scene sells them.
WHAT TO LEARN
Free-text emotion tags, an AI auto-tagger, and cloned character voices only read as real performance once picture, sound effects, and a consistent character voice carry the line with them.
Typing your own free-text emotional direction in brackets, like [delighted] or [scared, whispering], works even when the exact emotion isn't in the platform's preset tag list.
The same script, voice, and character can sound like three different performances purely from changing the bracketed tag in front of the line, so tag choice matters as much as voice choice.
An AI auto-tag suggestion feature can read a line's context and add tags, pauses, and emphasis automatically, giving a usable first pass before manual tuning.
Small wording changes toward how people actually talk, like a contraction instead of the formal phrase, measurably improved how natural a generated line sounded, independent of any emotion tag.
Stacking two tags, like [nervous, frustrated], reads as a different, more specific performance than either tag applied alone, the way a director layers notes for a real actor.
Cloning a distinct voice for a recurring character, instead of picking the closest match from a stock library, keeps dialect, pacing, and rhythm consistent every time that character speaks.
Building multiple emotional variants of one cloned voice, like a calm version and a panicked version of the same captain, lets a single character shift states while still sounding like themselves.
Judging an AI voice on the audio alone is misleading; pairing the same line with picture and sound effects is what actually determines whether a scene feels real.
Glossary
Terms worth knowing.
Fish Audio
An AI text-to-speech and voice-cloning platform. This video demonstrates its emotion-tag system, Story Studio script editor, and Instant Voice Clone feature.
Emotion tags
Bracketed natural-language directions typed in front of a line, like [delighted] or [scared, whispering], that tell Fish Audio's model how to perform that line's emotion.
Auto Tag
A Fish Audio feature that reads a line's context and automatically inserts suggested emotion tags, pauses, and emphasis, giving a usable first pass without manual tagging.
Story Studio
Fish Audio's multi-line script editor with a timeline, letting you assign voices per line, adjust timing, and add sound effects or music to build a full audio production.
Instant Voice Clone
A Fish Audio feature that turns a short recorded sample into a reusable custom voice, used here to build character voices that don't exist in the stock library.
Stable of characters
Multiple cloned voice variants of the same character, such as a calm version and a panicked version of one captain, that share consistent dialect and pacing across emotional states.
“So there you have the exact same script, the same voice, the same character, but three different emotional arcs.”
sets up the whole video's thesis in one sentence→ TikTok hook↗ Tweet quote
02:50
“It's just like you're talking to a voice actor in a studio.”
reframes prompting an AI voice as directing a real performer→ IG reel cold open↗ Tweet quote
07:40
“This is an important concept to embrace because you can create a stable of characters who each have different nuances.”
names the reusable production concept in one line→ newsletter pull-quote↗ Tweet quote
08:40
“It is the video combined with the audio and any sound effects which makes the whole experience.”
the video's real argument about how to judge AI voice tools→ newsletter pull-quote↗ Tweet quote
The Script
Word for word.
Read-along
Don't just watch it. Burn it in.
See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.
17px
metaphoranalogy
I can't believe you did that. I
can't believe you did that. I can't believe you did that. So there you have the exact same script, the same voice, the same character, but three different emotional arcs.
And the only thing that changed between those three was what I typed into these little brackets. This is part of fish audio's emotional control system which allows you to more accurately direct the performances of all of your voices by using natural language to direct the emotions that are used during the performance. But this video isn't just about adding tags to make something sound happy or angry or sad.
It's about what is the practical application of these tags. How much control over a character's performance do we actually have? And, most importantly, how does the expression of that realistic emotion impact the effectiveness of what you're delivering?
So can we actually create things like comedy and tension and fear or suspense or have characters lose emotional control over the course of time? And can we carry all of this into full conversations between characters?
That's what we're gonna find out today. You have a couple of places in the fish audio platform where you can use these emotion tags. First of all, is in the regular text to speech section.
So, if you just have something like I can't believe you did that with no emotion tags whatsoever and generate that speech in just a couple of seconds, you're going to have this. I can't believe you did that. Fish actually gives you two different versions to choose from.
I can't believe you did that. So right there, even without any emotion tags, you've got a really nice realistic read. But what if that's not the particular emotion you want expressed?
Or if we want to add some specific emotion to that, we've got several ways we can do it. One easy way is just to put the cursor in front of what you want to impact and click on tags and you can choose from any of these here. But you're not limited to just these choices.
You're encouraged to use your own natural language and description of the emotion you want portrayed when you're typing inside these tags. So, you'll notice that delighted is not one of these tags. However, if I go ahead and type delighted at the front of this line and generate this, you're gonna get two different interpretations of delighted.
I can't believe you did that. I can't believe you did that. Or if you change the word to disgusted.
I can't believe you did that. I can't believe you did that. Or combine multiple tags like here with scared and whispering.
I can't believe you did that.
I can't believe you did that. So let's choose a different voice, and I'll give you some other examples. We can add voices right from the text box here by clicking on more voices or over here under settings.
We can also click this icon here, and we have access to all the choices platform that we can search based on what we're looking for. We also have our voices that we've created on our own. So let's use this British voice and say a simple tag like nervous.
I don't even know what I'd say to her. I just make a fool of myself.
Myself. I don't even know what I'd say to her. I just make a fool of myself.
But what if we took that nervous tag and turned it into nervous and frustrated
like he's failed many times before and is getting angry. I don't even know what I'd say to her. I just make a fool of myself.
I don't even know what I'd say to her. I just make a fool of myself. So think about the possibilities here.
It's just like you're talking to a voice actor in a studio. You don't just have to say nervous, sad.
You can say, no. She's really broken your heart. I wanna feel that.
And if you look at some of these tags, we can also add effects like laughing and chuckling and moaning and sobbing and things like that. We've got several examples of all this coming up. Now, you may have noticed this auto tag button.
Well, what if you don't exactly know how you want to use this emotion tags or you're just learning or you'd like some suggestions from the AI. So first, let's hear this line with no emotion tags whatsoever. Oh my god, girl.
You have got to subscribe to Bob Doyle media. I'm not even kidding. You like that weird creative stuff.
Right? So it's a little flat. So let's click on auto tag and let the AI read the context of the line and suggest tags.
So you see it's got an excited tag. It's added a short pause. It's added some emphasis and another pause.
And let's see what difference that makes. Oh my god, girl. You have got to subscribe to Bob Doyle media.
I'm not even kidding. You like that weird creative stuff. Right?
So it's so much more realistic. Now let me jump in here with a little pro tip on making your voices sound even more realistic. Aside from the quality itself, the words that you say are very important.
As an example, in this script, it currently says, you have got to subscribe to Bob Doha Media. And that's fine. But most likely, a real person is gonna say, you've got to subscribe, not you have.
So let's just see what a difference that one little change made. Oh my god, girl. You've got to subscribe to Bob Doyle media.
I'm not even kidding. You like that weird creative stuff. Right?
Let's customize this even more by putting a giggle right in front of you like that weird creative stuff. So I'm just gonna end brackets, right, giggle, and regenerate. Oh my god, girl.
You've gotta subscribe to Bob Doyle Media. I'm not even kidding. You like that weird creative stuff.
Right? Okay. Let's go past just individual lines and let's put this into some sort of actual scene where the character's emotion changes over the course of the scene.
To do this, I'm gonna take advantage of the story studio inside of fish audio. This allows me to paste an entire script down the line in one shot and then go through and change voices and make all of my changes. Here you can see that I've got a script that's already done.
And, another advantage of using story studio is that we have a timeline here where we can change the timing of where the various voices land. We can add sound effects and music to make it an entire audio production. So before I play this, let me just kind of walk you through the arc here.
Now, it's a fake cooking show, and this guy is demonstrating something seemingly simple, but something goes weird. His saucepan starts making a noise, and the sauce inside the pan starts talking to him and he is now reacting in front of a live audience. So here's how this plays out.
And now we add just one teaspoon
of chili powder. That seemed like significantly more than a teaspoon. Why is the saucepan making that noise?
Feed me. That's that's perfectly normal.
Some sauces do become sentient.
Feed me now.
We'll be right back after this message.
So now we can export this entire thing and it's already basically produced. We could use it for the background of a video or if it's an audio file just use it on its own.
But now, sometimes the generations are gonna come out and you're gonna wanna change it. The beauty of AI is that it's very easy to iterate. If you don't like the output of a particular line, you can just keep regenerating it.
For example, in this line here, why is the saucepan making that noise? It sounds like this. Why is the saucepan make in that noise?
So, it says saucepan first of all. And, I would definitely change that to saucepan and then click regenerate.
Why is the saucepan making that noise? And, if I didn't like that exact read, I could generate it again. Why is the sauce pan making that noise?
Again, just like being in the studio with a real voice actor. We're not just choosing random emotions because we can. We're trying to walk you through the natural emotional arc this character would have if he was doing a cooking show and his sauce started talking to him while he tried to maintain his professionalism.
Story studio is also the natural place to create dialogue between characters. And I got some really fun examples of characters interacting with all of these emotional tag put in to make it seem much more real. But speaking of realism, I got another pro tip for you in terms of making the character sound exactly like you want.
In the various scenes I'm gonna show you, I created some custom characters because I wanted them to read a particular way. For example, I wanted one character who sounded kinda like a really confident airline pilot. So instead of searching the library for one and trying to force it in there, I knew what I wanted it to sound like.
So I just went to create a voice and I went to instant voice clone and recorded exactly what I wanted. Ladies and gentlemen, everything is going great and we should be within about thirty minutes or so. That ends up creating a voice that sounds like this.
All systems are looking good and we're ready for a wonderful journey. And, for another character, I cloned my voice recording something like this so that my character would sound like this. I'm a little nervous about this journey, sir.
I didn't expect to be here. This is an important concept to embrace because you can create a stable of characters who each have different nuances. It could be the calm version of the captain or it could be the we're in a panic mode and then using the correct version with the correct script in conjunction with the emotion tag and now you've got ultimate control over the realism.
Every time you use that voice it'll have a consistent dialect. It'll have the same pacing and overall rhythm. So that every time you use it, it sounds like the same person.
So now we have all the pieces we need. We got the studio to create this in. We have a stable of voices, some of which we've created, some of which are in the platform already, and we have full control over the emotional expression of all of our characters.
But just before we get into these examples, there's one other thing I want you to take into consideration. Consideration. When When we we just listened to that example of the guy with the cooking show, there was no video going with it.
I was asking you to use your imagination to fill in the blanks of what that might look like. And, because we're evaluating a text to speech platform, you're probably fixated on does that sound real? Does everything sound exactly real?
And that's not the way that we normally consume this type of thing, especially if we're watching videos. It is the video combined with the audio and any sound effects which makes the whole experience. So for all of these examples, I'm taking the output from fish audio and we put video to them so that you can see them used in proper context.
Because I think most people watching my channel are making videos, and they want them to occur as realistically as they want them to occur. And using techniques like this can be your secret weapon. Why do you know so much about this?
Because I actually trained for knighthood.
Could you perhaps train me very quickly?
The battle starts in six minutes. Excellent. Intensive course.
Fine. Put your left foot in the stirrup.
He's looking at me. It's a horse. He knows.
Knows what?
That I'm weak.
Get on the horse.
It's moving. You haven't even mounted it yet.
Then why am I falling?
Who is she? Before you get upset. That's never been a promising opening.
She's an AI version of you.
An AI version of me. Yes. I trained her on your voice, your mannerisms, your preferences.
My preferences? Just the useful ones. What does that mean?
She likes documentaries.
I like documentaries. You complain during documentaries. Because they leave things out.
See? She doesn't do that. Welcome back number two.
You really went all out. At last, a headquarters worthy of my genius. Who are those people in the kitchen?
What people? The family making pancakes. Why is there a family making cakes in my secret lair?
Funny story. I specifically requested an unfunny story.
There was a small cash flow issue during construction. And? So I listed the East Wing online.
You listed my evil headquarters on Airbnb?
Only weekends.
This is a secret lair. All systems are nominal. Course is locked.
We are officially on our way to Mars. Great. Then I can finally stop pretending I understand any of these buttons.
That's why they sent me. Quick question. Go ahead.
Where do we keep the extra oxygen tanks? Storage Compartment 3. Storage Compartment 3 is empty.
No. It isn't. I'm looking directly at it.
So emotion tags aren't just cool because they can make a person sound happy or angry or sad. It's because now you can tell a real story with real characters with believable emotions. And when you build a stable of characters like I talked about before, you have a whole new level of consistency as well.
I've got all the information linking to fish audio down in the description. If these are the types of technologies and techniques you like to learn about, I invite you to subscribe to this channel. Subscribe to Bob Doyle Media.
Because this is the type of stuff we play with all the time. And now we add just one teaspoon of chili powder.
That seemed like significantly more than a teaspoon. Why is the saucepan making that noise? Feed me.
That's that's perfectly normal. Some sauces do become sentient.
Feed me now. We'll
be right back after this message. Today,
we ride into battle and defend the kingdom. For honor,
for glory, for the king.
Mount your horse, sir Edmund. Right.
Go ahead. I am. You are standing next to it.
I know where the horse is.
Have you ever ridden a horse? Of course, I've ridden a horse. When?
Childhood. How old were you? Four.
That doesn't count. It moved.
Your mother was holding you.
Why do you know so much about this? Because I actually train for knighthood. Could you perhaps train me very quickly?
The battle starts in six minutes. Excellent. Intensive course.
Fine. Put your left foot in the stirrup.
He's looking at me. It's a horse. He knows.
Knows what?
That I'm weak. Get on the horse.
It's moving. You haven't even mounted it yet.
Then why am I falling? Four. I tripped over the bucket.
The kingdom is doomed.
Could we defend it on foot? Against cavalry? Then today, we invent tactical jogging.
Who is she? Before you get upset. That's never been a promising opening.
She's an AI version of you.
An AI version of me. Yes. I trained her on your voice, your mannerisms, your preferences.
My preferences?
Just the useful ones. What does that mean?
She likes documentaries.
I like documentaries. You complain during documentaries. Because they leave things out.
See? She doesn't do that.
Oh, I see. So she's better than you. Not better.
Smarter?
Different. Prettier? That's not even a setting.
Then why is she wearing my sweater?
She picked it.
You let a computer wear my sweater.
Okay. I realize we're no longer discussing the technical achievement.
You built someone who sounds like me, acts like me, and apparently annoys you less.
That's not why I made it. Then why? Because you've been traveling so much and the house felt empty.
That is either incredibly sweet or deeply alarming.
Can it be both?
Probably. So where is she now?
Standing behind you. What?
Hi. I reorganized your closet.
Delete her. Welcome back number two. The new secret lair is finally operational.
You really went all out. At last, a headquarters worthy of my genius. Who are those people in the kitchen?
What people? The family making pancakes. Why is there a family making pancakes in my secret lair?
Funny story. I specifically requested an unfunny story.
There was a small cash flow issue during construction. And? So I listed the East Wing on online.
You listed my evil headquarters on Air B and B? Only weekends. This is a secret lair.
Do you have any idea what Waterfront Volcano property costs?
They are taking selfies next to the death ray.
I did ask them not to touch it.
There's a child sitting in my command chair.
He gave us five stars.
I don't care about the rating.
We were at 4.6 before him.
What did he say in the review?
Amazing views, fun Shark Tank. Host seemed emotionally unavailable.
Emotionally unavailable?
You do tend to brood.
I am a supervillain.
Right. And check out is at eleven.
Fine.
But nobody uses the doomsday device before breakfast. All systems are nominal. Course is locked.
We are officially on our way to Mars. Great. Then I can finally stop pretending understand any of these buttons.
That's why they sent me. Quick question. Go ahead.
Where do we keep the extra oxygen tanks? Storage Compartment 3. Storage Compartment 3 is empty.
No. It isn't. I'm looking directly at it.
Then you're looking at the wrong compartment. There's only one compartment with a giant label that says oxygen. What does ah mean?
It means I may have removed them before launch. You removed the oxygen? They were heavy.
They're oxygen. I was trying to improve fuel efficiency by making the astronauts lighter.
When you say it like that, it does sound poorly conceived. How much oxygen do we have left? Enough?
Enough for what? A very meaningful conversation. Turn this ship around.
Already doing it. Can we make it? Absolutely.
Probably. Those are extremely different answers. Look, worst case, we're about to become a very inspirational documentary.
The Hook
The bait, then the rug-pull.
Bob Doyle opens with the same four words spoken three different ways, then spends eighteen minutes proving the difference isn't a gimmick — it's a production tool that can carry a script, a character, and a whole sketch.
Frameworks
Named ideas worth stealing.
01:35concept
Natural-Language Emotion Tags
Type any free-text emotional direction in brackets in front of a line, not limited to the platform's preset tag list, and the model performs it — e.g. typing [delighted] even though 'delighted' isn't one of the built-in choices.
Steal forany TTS-based voiceover or character line that needs a specific, non-generic emotional read
03:13concept
Auto Tag
An AI suggestion feature reads a line's context and inserts its own tags, pauses, and emphasis automatically, turning a flat read into a directed one without the creator having to know which tag to use.
Steal fora fast first-pass on any generated line before manual tag tuning
03:50concept
Contraction Pro-Tip
Swapping formal phrasing ('you have got to') for the contraction real people actually say ('you've got to') measurably improved how natural a generated line sounded, independent of the emotion tag applied.
Steal forany TTS script — write it the way people actually talk, not grammatically formal
07:40concept
Stable of Characters
Clone multiple emotional variants of the same voice, such as a calm captain and a panicked captain, so a recurring character stays consistent in dialect and pacing across every scene while still shifting emotional state.
Steal forany recurring AI character used across multiple videos or an episodic series
CTA Breakdown
How they asked for the click.
VERBAL ASK
11:14subscribe
“If these are the types of technologies and techniques you like to learn about, I invite you to subscribe to this channel. Subscribe to Bob Doyle Media.”
Soft CTA placed right after the thesis wrap-up line and right before the full five-sketch reel plays back to back, so the ask lands before the payoff instead of interrupting the walkthrough.
Add Modern Creator as a preferred source and Google shows you more of our breakdowns in Search, Top Stories, and AI Overviews. It only changes what you see, and you can undo it in your Google settings anytime.
Add to Preferred SourcesOpens your Google source preferences with us pre-loaded. Tick the box and you're done.
Neel Dhingra sits down with Think Media CEO Sean Cannell to break down the one long-form video a business owner can make once a week that teaches, sells, and closes clients for you.
A start-to-finish OBS build for horizontal and vertical multistreaming — scenes, overlays, split audio, alerts, and the actual bitrate/keyframe numbers to use.
A nine-minute, six-step walkthrough for shooting a talking-head YouTube video on a smartphone — topic, background, gear, recording, review, and editing.