Skip to main content

Try it live in the API playground

Drop text with the markers below into the text field and send a real request to hear the emotion.

Overview

Fish Audio models support 64+ emotional expressions and voice styles that can be controlled through text markers in your input. Add natural pauses, laughter, and other human-like elements to make speech more engaging and realistic.
This page shows S2 usage with [bracket] cues. If you use the legacy S1 model, wrap markers in parentheses instead — see S1 (legacy) syntax below for the full list, or the Models Overview.

How It Works

Add emotional or stylistic cues in square brackets within your text:
The S2 TTS models will interpret these markers and adjust the voice accordingly.

Complete Emotion Reference

Basic Emotions (24 expressions)

Advanced Emotions (25 expressions)

Sound & Delivery Markers

These markers aren’t emotions — they shape how a line is delivered, add natural human sounds, or layer in ambient effects. Combine them with the emotion cues above.

Tone Markers (6 expressions)

Control volume, intensity, and emphasis. Place [emphasis] right before the word or phrase you want to stress:

Audio Effects (11 expressions)

Add natural human sounds:

Special Effects

Additional markers for atmosphere and context: You can also use natural expressions like “Ha,ha,ha” for laughter without tags.

Usage Guidelines

Placement Rules

For S2:
  • Sentence-level emotion cues usually work best at the beginning of sentences
  • Tone controls can go anywhere in the text
  • Sound effects can go anywhere in the text
  • Bracket cues can use natural language descriptions and are not limited to a fixed set of tags
Correct:

Advanced Techniques

Combining Effects

You can layer multiple emotions for complex expressions:

Emotion Transitions

Create natural emotional progressions:

Background Effects

Add atmospheric sounds:

Intensity Modifiers

Fine-tune emotional intensity with descriptive modifiers:

Language Support

All 13 supported languages can use emotion markers. For sentence-level control, cues usually work best at the sentence start in these languages:
  • English, Chinese, Japanese, German, French, Spanish, Korean, Arabic, Russian, Dutch, Italian, Polish, Portuguese

Best Practices

Do’s

  • Use one primary emotion per sentence
  • Test different emotion combinations
  • Match emotions to context logically
  • Add appropriate text after sound effects (e.g., “Ha ha” after laughing)
  • Use natural expressions when possible
  • Space out emotional changes for realism

Don’ts

  • Don’t overuse emotion tags in short text
  • Don’t mix conflicting emotions
  • Don’t make bracket descriptions so long that they interrupt readability
  • Don’t forget brackets
  • Don’t place sentence-level emotion cues far from the sentence they control

Common Use Cases

Customer Service

Storytelling

Educational Content

Marketing & Sales

Troubleshooting

Emotion Not Working?

  1. Check placement - Put the cue where the emotion or effect should begin
  2. Keep wording clear - Use concise natural language descriptions
  3. Use the right syntax - S2 cues use square brackets; S1 cues must use parentheses

Unnatural Sound?

  • Space out emotional changes
  • Use appropriate intensity
  • Test with different voices
  • Add context text after sound effects

Performance Notes

  • Emotion markers don’t count toward token limits
  • No additional latency for emotion processing
  • All emotions available on all pricing tiers
  • Maximum of 3 combined emotions per sentence recommended

Quick Reference Tables

Emotion Intensity Scale

Common Combinations

S1 (legacy) syntax

The default S2-Pro model uses [bracket] cues with free-form natural language. The previous-generation S1 model uses the same emotion names but requires (parentheses) and a fixed tag set:

See Also