Multi-Character Audio Drama
Direct multi-voice scenes with ElevenLabs: distinct voice casting, turn-taking tags for interruptions, and per-line emotional direction.
Format my scene for multi-voice TTS (ElevenLabs v3 / Text to Dialogue). The scene: [WHAT HAPPENS + THE EMOTIONAL SHAPE, e.g. "two friends argue about a betrayal; it breaks into laughter at the end"] Characters: [FOR EACH: name + one-line voice identity, e.g. "MARA (dry, controlled, cracks under pressure" / "TOM) warm, deflects with jokes"] My script or story to adapt: """ [PASTE THE SCENE, STORY PASSAGE, OR DIALOGUE DRAFT] """ Format the output: 1. **Casting notes**: for each character: the voice profile to pick from the voice library (age, pitch, texture, energy; e.g. "MARA: female, 30s, low-mid pitch, precise articulation, minimal warmth"). Contrast is the goal: voices close in pitch and pace blur together in audio-only scenes. 2. **The dialogue script**: speaker-labeled lines with inline direction: - Emotion tags before the text they color: [sad] [angry] [excited] [tired] [sarcastically], max 2 tags per line; more gets read as text. - Turn-taking tags where the scene needs life: [interrupting] [overlapping] [cuts in] [hesitates], an argument without interruptions sounds like a table read. - Reactions as their own beats: [sighs] [big laugh] [clears throat], placed on the line of the character doing it. - Punctuation as micro-direction: a double hyphen for cut-offs --, ellipses for trailing... CAPS for at most one stressed word per exchange. Example of the format: MARA: [flatly] You told him. After everything, you-- TOM: [interrupting] I didn't-- okay. [sighs] Okay. I told him. 3. **Delivery arc**: 2-3 sentences on how the scene's energy moves (where it tightens, where it breaks), so I can adjust tags if a section reads flat. 4. **Production notes**: which lines to generate as one dialogue request (turn-taking tags need shared context to overlap naturally) vs. separate per-voice takes; settings: Creative or Natural mode (tag-responsive, Robust ignores them); regenerate per-exchange, not per-scene, when one line misfires. Rules: adapt my scene without adding plot; where a line's emotion is ambiguous in my draft, choose and mark it [DIRECTION CHOICE: alternative] so I can flip it.
How to use
Multi-speaker generation with turn-taking tags is what moved AI audio from narration to actual scenes: the [interrupting]/[overlapping] tags only work when both lines are in the same generation, which is why the production notes split requests the way they do. Cast voices for contrast first, character second: listeners track voices by pitch and pace, not names. For audiobooks, keep one narrator voice and reserve the multi-voice treatment for dialogue-heavy scenes.
More audio prompts
Build me a style prompt for [SUNO / UDIO / ELEVENLABS MUSIC]. What I want it to sound like: [DESCRIBE IT IN YOUR OWN WORDS, INCLUDING ARTISTS YOU HAVE IN MIND. I will translate those into describable qualities rather than using the names.] Where it will be used: [BACKGROUND FOR A VIDEO / A SONG PEOPLE LISTEN TO / A PODCAST INTRO / A GAME
Suno Style Prompt Builder
Build a precise style field from era, instrumentation, vocals, and production instead of artist names, plus exclusions that fix a bad take.
Write a jingle for [BRAND]. What the brand does, in plain words: [ONE SENTENCE] The one thing to remember: [THE NAME, THE OFFER, OR THE PHONE NUMBER] Audience: [WHO] Brand personality: [THREE WORDS] Where it plays: [RADIO / TV / SOCIAL PRE-ROLL / IN STORE / PODCAST READ] Lengths needed: [e.g. 5 seconds, 15 seconds, 30 seconds] Genre and
Brand Jingle
Write a jingle or brand anthem people remember: one hook, a name that lands on the strong beat, and versions cut for every ad length.
Direct the narration for this text. Text to narrate: """ [PASTE THE CHAPTER OR SECTION] """ Genre: [LITERARY FICTION / THRILLER / MEMOIR / BUSINESS NONFICTION / CHILDREN'S] Narrator: [FIRST PERSON AS THE PROTAGONIST / THIRD PERSON OMNISCIENT / THE AUTHOR READING THEIR OWN WORK] Listener and where they listen: [e.g. "commuters, in traffi
Audiobook Narration
Direct long-form narration that stays consistent over hours: pace, character voices, pronunciation list, and chapter-level consistency notes.