← Blog

How-to seriesSeptember 28, 2026

How to create a voiceover for a script

Turn an approved script into a voiceover with Vaaya, choose a supported voice, and check pronunciation, pace and the downloaded audio before using it.

A voiceover is ready when someone has listened to the file and checked the words. A successful generation response does not tell you whether a product name sounds right or whether the last sentence fits your video.

This walkthrough uses Vaaya to create one narrated file from an approved script. It includes a short review step before any revision. The script and handoff example below are illustrative; no audio was generated for this tutorial.

1. Prepare a script someone can say aloud

Connect your agent using the installation guide, choose where the audio will be used and set a total budget. Use text you are entitled to narrate and a voice you are permitted to use.

Write the spoken version of numbers, abbreviations and names. A line such as “Check v2.5 at 09:30” may need to become “Check version two point five at nine thirty.” Read the script aloud once. Long sentences that are awkward for you are likely to need editing after synthesis too.

Keep screen labels and production instructions outside the text to be spoken. If a pronunciation matters, provide the intended reading and ask whether the selected model supports a pronunciation control. Do not assume it accepts another provider's markup.

2. Choose the voice and price the exact script

The current Vaaya catalog includes fal/generate with model elevenlabs--tts--turbo-v2-5. This route accepts text, a supported preset voice name and speed. Vaaya's registered default speed is 1.0; the model supports values from 0.7 to 1.2.

Ask the agent to check the current voice options and model schema before choosing. A provider voice identifier from a different endpoint is not automatically a valid preset here. Quote the exact script and settings you plan to submit. Speech pricing can depend on text length, so a short sample quote is not a quote for an entire chapter.

Use this prompt:

Use Vaaya to narrate the approved script below.
Purpose: [demo, training clip or other destination].
Voice: [supported preset], with [desired delivery].
Pronunciation notes: [names and intended readings].
My total budget, including revisions, is [amount].
Check the current model schema, voice and quote before generation.
Keep every call within the remaining budget. Generate one take,
retrieve that job, and return the audio file with the exact script,
voice, settings and actual charge.
Script: [paste only the words to be spoken]

3. Generate once and retrieve the file

Have the agent pass the script through the model's text field, with an explicit voice and supported speed. Apply a max_cost_cents ceiling to the call. If Vaaya returns a job ID, use result to check that job instead of purchasing another take.

Download the audio when it is ready. Keep the original returned file before converting it for an editor or presentation. Open it and check its measured duration and format; neither the prompt nor a filename proves those properties. See the Vaaya tool reference for the generation and result flow.

4. Listen for the things a status field cannot check

Listen from start to finish while following the approved script. Check that every sentence is present, names are understandable and pauses occur where you intended. Then listen once without reading. That makes rushed transitions and strange emphasis easier to catch.

For a video, place the narration against the actual visuals. If the spoken instruction arrives after the action on screen, adjust the edit or script before making the delivery faster. A text-to-speech model does not know the timing of an unseen screen recording.

This is an illustrative handoff format:

Field Example or value to record
Script version demo-intro-v1.txt
Audio file demo-intro-v1.wav
Measured duration Fill from the downloaded file
Pronunciation review Record the words actually heard
Approved for use Yes or no, after listening
Actual charge Copy from the call record

5. Fix the script before buying repeated takes

If a term is mispronounced, revise its spoken spelling or use a supported pronunciation feature. If the recording is too long, remove words before increasing speed. Preserve the first file so you can compare the revision.

An empty result or failed job needs error handling, not an assumption that narration finished. Check the returned error, input limits and voice selection. A running job needs another status check. A completed take with poor delivery needs a deliberate revision and a new budgeted call.

Store the accepted audio beside its exact script, voice, settings and receipt, so the next section can use the same delivery choices.

Questions

Which Vaaya route can create this voiceover?

The current catalog includes fal/generate with the elevenlabs--tts--turbo-v2-5 model key. Check the current schema, permitted voice options and quote before running it.

Can I specify the exact audio duration?

The duration depends on the script, voice and delivery. Measure the finished file and revise the text or supported speed setting if it needs to fit a fixed slot.

Does a transcript prove that the voiceover is correct?

No. A transcript can help find missing words, but listening is needed to check pronunciation, pacing, emphasis, clipping and pauses.

Try Vaaya with your agent.

Connect your agent, choose a service, and try your first call with Vaaya.

npx @vaaya/mcp install