Home
Assets
My Works
Voice Library
Developer Docs
Blog
English
Pricing
Sign In
Multi-Speaker Voiceover
HomeAudio ToolsMulti-Speaker Voiceover

Audio Tools · Dialogue Scripts · 12 Languages

Multi-Speaker Voiceover

Paste a dialogue script, give every speaker their own voice

Paste a script, or write it one segment at a time.
HOW IT WORKS

How multi-speaker text to speech works

Paste a script and it splits into lines, then gives each speaker name its own voice. Every line has a role dropdown, so you can move any line to a different speaker.

  1. 01

    Paste your dialogue script

    Paste the whole script at once. Each block of dialogue becomes its own segment — split on blank lines, or on line breaks when there are no blank lines — and nothing gets dropped. You can also start from an empty segment and write it here one block at a time.

  2. 02

    Let it sort out the speakers

    Names written as "Alice:" or [Alice] at the start of a paragraph are picked up and turned into roles. If the script has no name prefixes, every segment lands on role 1 and you assign the speakers yourself.

  3. 03

    Give each role a voice

    Open the voice picker for each role and choose a voice, plus an emotion if you want one. A name that matches a voice in the library is bound to that voice already; the other names get voices in the order they first appear.

  4. 04

    Adjust the read, then export

    Use the role dropdown on any segment to move a line to a different speaker, and insert a pause or a laugh where the read needs one. Then generate and export the audio with SRT subtitles.

Q&A

Multi-Speaker Text to Speech FAQ

01

What is multi-speaker text to speech?

It reads a dialogue script aloud with a different voice for each speaker, so one script becomes a conversation. You paste the script, assign a voice per role, and export a single audio file.

02

How does it know who is speaking?

It looks for a speaker name at the start of each paragraph — "Alice:" with a half-width or full-width colon, or [Alice] in brackets — and names can be up to 12 characters. The speaker split is applied only when at least 60% of paragraphs carry a name and at least two different names appear; otherwise everything goes to one role and you assign the speakers yourself.

03

How many speakers and how much text can I use?

Up to 10 roles, 200 segments, and 20,000 characters in total. You can generate in 12 output languages.

04

Can I change who says a line after the split?

Yes. Every segment has a role dropdown, so you can move any line to a different speaker at any point. Right after a paste an undo also appears once — it can revert the paste, or just undo the speaker split and keep the text.

05

Can I control how a line is read?

Yes, in two ways. Each voice can carry an emotion, set inside the voice picker; the default is automatic, where the read is performed from the text itself. You can also insert tone tags at the cursor: laugh, chuckle, sigh, inhale, throat-clear, hesitation, and pauses of 0.5s, 1s, or 2s. Tag support depends on the voice — pauses work with more voices than the onomatopoeia tags, which are not available on Japanese voices, and a tag the selected voice cannot handle is flagged in the editor and blocks submission until you remove it.

06

Can I download the audio?

Yes. Export the finished audio file, and export SRT subtitles alongside it when you need captions.

Where multi-speaker voiceover fits
07

Two-host podcast scripts

You wrote the episode as a back-and-forth between a host and a guest. Give each one a voice and both parts come out sounding like two people.

08

Language lesson dialogues

Voice a practice conversation so the learner hears two distinct speakers. Export the SRT alongside the audio for on-screen text.

09

Audio drama and fiction

Give each character in a scene its own voice and emotion, then drop in a sigh or a one-second pause where the writing needs a beat.

10

Ad scripts and skits

Short two-character scripts for a promo or a social clip. Write it in the editor, hear it read back, and adjust the roles until the exchange lands.

ListenHub

One platform for AI content creation — from an idea to audio, images, slides and video.

Audio
AI PodcastText to SpeechMulti-Speaker VoiceoverAudio to TextVoice CloningAI Voice
Video
AI VideoExplainer VideoLip SyncAd GeneratorPromo VideoMotion Transfer
Image
AI ImageSlides
Resources
PricingDiscoverVoice LibraryBlogAPI DocumentationMCP Server GuideAgent Skills Guide
Company
About UsContact UsTerms of UsePrivacy Policy
© 2026 MarsWave