Back to all models

ElevenLabs / Voice

ElevenLabs.
Give the story its voice.

ElevenLabs provides speech generation models that turn text into spoken audio. It belongs in the sound layer of an AI video workflow: narration, character dialogue and local-language tracks. The voice needs to carry the same intention as the scene it accompanies.

Inside the storyVoice
Fictional characters Anaya and Kabir facing each other across a warmly lit table.
AI Director’s shot briefThe same intention, in another language.
01 / Script

Approve the meaning, register and character names.

02 / Voice

Audition the selected voice in the intended language.

03 / Timing

Fit the pause and line length to the scene’s edit.

Fictional story concept. Illustrative artwork, not output from the featured model.

Model capabilities

What ElevenLabs brings to the soundtrack

Text to speech
Generate spoken audio from a script using a selected voice.
Multilingual speech
Supported models cover multiple languages. Language availability and the suitability of a voice vary by model and track.
Voice and delivery choices
Voice selection and model-specific controls help shape delivery. Audition the actual dialogue rather than relying on a generic sample.

Capabilities describe the model families, not a guarantee that every version or control is available in Tosheo. Production access depends on provider availability, references, rights and the approved scope.

02 / ElevenLabs

Where ElevenLabs fits in your story

ElevenLabs fits narration, dialogue drafts and language-track production. Tosheo’s AI Director can connect the script to a character voice, pronunciation guidance and the relevant edit. Local-language review checks meaning and register as well as the sound of a line, so translation remains a storytelling decision.

Fictional exampleThe same intention, in another language.

Kabir says Anaya’s name quietly. Keep the hesitation and the relationship intact when adapting the line.

Script
Approve the meaning, register and character names.
Voice
Audition the selected voice in the intended language.
Timing
Fit the pause and line length to the scene’s edit.

03 / ElevenLabs

Plan the ElevenLabs handoff

The AI Director prepares the voice request from the scene plan, using AI Skills to specify the framing, performance or sound intention. You can review creative preparation or use Auto mode within your chosen scope. The model receives a bounded job, with the references and constraints that matter to that asset.

For an independent creator, this connects a creative decision to the next production step. For a publisher, it keeps recurring assets attached to the episode plan. Brand buyers can carry approved product references into the shot; production houses can retain the treatment and handoff requirements across the sequence.

04 / ElevenLabs

One identity, several languages

The approved voice reference is canon in the same way a face is. It travels with the character, it is inherited by every shot, and changing it requires a review.

The part that is specific to voice is that the identity has to survive a change of language. This is why language tracks are declared at greenlight: the voice is chosen against every declared track before it is approved, rather than a second voice being sourced afterwards to match a performance it was never cast for.

  1. 01
    Declare the tracks at greenlight

    Define the primary language and any additional tracks at greenlight. Confirm the selected voice works in each before approving it.

  2. 02
    Cast the voice against all of them

    An identity that only holds in one language has failed before production starts.

  3. 03
    Record pronunciation as canon

    Names, places, regional phrasing and register. Checked in QA rather than re-decided per session.

  4. 04
    Review each track natively

    A named approver per track. A Hindi reviewer cannot sign for Tamil.

  5. 05
    Repair at segment level

    A mispronounced line is an audio segment, not a track and not an episode.

05 / ElevenLabs

The permission rules, stated plainly

  • No cloned voice of a real person without verified permission
  • No celebrity likeness or voice, in any genre, without verified permission
  • No voice used outside the territory, term or platform its grant covers
  • No deceptive testimonial production
  • No removal of synthetic-media disclosure where it is required

Where a permission exists, it is recorded with its territory, term, platform and expiry like any other grant, and the expiry is a date the system holds rather than something a person is expected to remember two years later.

06 / ElevenLabs

Evaluating a voice model in the actual scene

Provider documentation

Capabilities checked against official sources. Story applications are Tosheo’s editorial examples, not comparative benchmarks.

ElevenLabs text-to-speech documentation

Questions before production

Can I clone my own voice?

With verified permission - which, for your own voice, means verifying it is yours - yes, and it is recorded as a grant with a territory, term and expiry. The verification step is not a formality we can skip because the person asking sounds confident.

How do you stop a character sounding different in another language?

By casting the voice against every declared language track before approving it, rather than matching a second voice to a finished first performance. It is the single decision that most determines whether a regional track works, and it is made at greenlight.

Who fixes a mispronounced name?

Pronunciation guidance is canon, so the fix is a canon correction plus a targeted repair of the affected audio segments. The picture, the other tracks and the rest of the episode are untouched.

Is synthetic-media disclosure applied to audio?

Disclosure obligations attach to the deliverable and are checked in the delivery package. For serialized production at launch, captions and synthetic-media disclosure are on by default.

Do you use one voice provider for everything?

Routing considers language coverage and pronunciation credibility per track, so a production can use different providers for different tracks. Whatever is used is recorded in the provenance for each accepted asset.

AI Director + AI Skills + Models

The model makes an asset. Your AI Director connects the story.

  1. 01

    Your direction

    Share the brief, format, episode plan and intended languages.

  2. 02

    AI Director

    Build the storyboard, references and shot-level direction with AI Skills.

  3. 03

    Model fit

    Match the asset to an available model and the approved production scope.

  4. 04

    Story-ready assets

    Check the result in context, repair what needs attention and assemble the edit.

Review creative preparation yourself or use Auto mode within your chosen scope. Cost authorisation, rights clearance and final delivery acceptance remain explicit decisions.

More from the model library

Keep building the story.

Tosheo / AI-native production

Start with the scene.
Let the direction connect it.

See how Tosheo brings the brief, AI Skills, generation and review into one production flow.