How to Make a Sales Voiceover with ElevenLabs
A practical 2026 path: write a 15–45 second sales or product-demo script, pick a Voice Library voice, generate the read in Text to Speech, tweak pacing, download WAV, and drop it on the demo or landing-page timeline. Built from ElevenLabs' live Text to Speech help and pricing (August 2026).
Most “AI sales voiceover” posts skip the two facts that actually decide the file. First: a thirty-second pitch is not a podcast sting, and it is not a talking-head demo. Second: a free-tier preview is not a track you put on a paid-media landing page. This guide is the working path for the spoken close — script, voice, generate, pace, download, drop it on the demo — using the one voice desk we send for this brief: ElevenLabs.
When a stock ElevenLabs voice is enough — and when you want a face or a booth
Use ElevenLabs when the open is a short, realistic read you will regenerate. A hero line on a pricing page. A product-demo bed under a screen grab. A sales-email clip that still sounds like the same person who closed last month’s offer. The economic case is the recut: edit the sentence, generate again, replace the clip. You do not re-book a booth for thirty seconds.
Skip ElevenLabs if the landing page needs a presenter walking the product on camera. That file is a talking-head demo, and the working path on this site is How to make product demo videos with Synthesia — an ordinary guide link, not a second money hop. Skip both if the missing piece is a directed studio narrator on slides. That is How to make training voiceovers with Murf. Murf is the foil, not a Try button. Skip a stock library voice if the brand is a named founder’s voice and you do not have a clone yet. That is How to clone your voice for YouTube.
The seven-step ElevenLabs sales voiceover
- 01
Write a 15–45 second sales script — problem, product, one proof, one CTA
A sales voiceover is a spoken pitch, not a show sting and not a two-minute explainer. Time it out loud: fifteen seconds is a hero-line plus a click; forty-five is a problem, the product name, one proof, and one next step. Open on the buyer’s cost of doing nothing, name the product, give one number or one outcome a prospect can check, then say the URL or the trial. Spell plan names, dollar figures, and abbreviations the way they should be heard on a landing page. If you wanted a 10–20s show open, that is the podcast-intro how-to. If you wanted a 2–5 minute YouTube narration, that is the YouTube-voiceover how-to.
- 02
Confirm a generated VO is the right sales format
Use ElevenLabs when the missing piece is a short, realistic read you can regenerate when the offer or the hero line changes. Skip it if the landing page needs a presenter on camera walking the product. That is a talking-head demo, and that how-to is an ordinary path on this site, not a hop here. Skip it if the file is a directed studio narrator timed to a slide lesson. That is Murf Studio, and that how-to is also an ordinary path. Skip a library voice if the brand is the founder’s voice and buyers would notice a stranger. That is a clone, and the full training-set path lives on the YouTube clone how-to.
- 03
Open Text to Speech — the speech playground, not Agents
ElevenLabs’ published product for this brief is Text to Speech: paste text, pick a voice, generate speech, download a file. That is the desk this page is written for. ElevenAgents is a different product (conversational agents). Studio is a longer project surface for chaptered work. Image & Video on the pricing page is a different surface — we are not inventing an MP4 download from the speech playground. This page sends you to ElevenLabs only.
- 04
Sit on a plan that grants commercial rights before you treat the file as the landing page
ElevenLabs’ own docs are explicit: you keep ownership of generated audio, but commercial usage rights come with paid plans. Free is an audition — published pricing lists 10,000 credits a month and no commercial license. Starter is the first paid tier that lists a commercial license and Instant Voice Cloning. Professional Voice Cloning is listed on Creator and above. Confirm the live grid on elevenlabs.io/pricing. We do not invent a checkout total or a commission rate. Do not publish a free-tier read on a paid-media landing page or a sales email.
- 05
Pick a Voice Library voice — clone only if the brand already is your voice
Open Voices and browse Default Voices or the Voice Library. Preview before you apply. Cast one narrator and keep it for the landing-page cut, the recut, and any matching demo bed. Filter toward narration-style reads, then listen on headphones — a sales VO should sound like a calm closer, not a cinematic whisper and not a training-module drone. A stock voice is enough when the brand has no existing vocal identity. Clone only if buyers would notice a stranger. Instant Voice Cloning is the published self-serve path from short samples (ElevenLabs’ cloning help: about 1–2 minutes of clean audio) on paid plans that list it. Professional Voice Cloning trains a dedicated model on a longer set (published as 30–180 minutes) and requires Creator or above. Do not invent a clone wizard here — if you need the training-set workflow, use the clone how-to. Do not clone a customer, a podcast guest, or a voice you found on someone else’s demo.
- 06
Paste the 15–45s script, pick a model, generate, then tweak pacing and re-roll
Type or paste the pitch into the text box as one short block. Voice first, then model, then settings — that is the order ElevenLabs’ Text to Speech guide ranks. For a short English sales read, Multilingual v2 is the published “most stable on long-form” model and still the safest default for a line you will reuse on a landing page. Eleven v3 is the expressive model; it supports audio tags such as [sighs] or [clears throat], and it does not expose every older slider. Official starting settings for the sliders that exist: Stability around 50, Similarity around 75, Style exaggeration at 0. Speed defaults to 1.0 (range 0.7–1.2) where the model offers it — nudge Speed, or add a dash / em-dash, when the close feels rushed. Spell out numbers. Then press Generate Speech. The model is nondeterministic — same text can yield a different take. Re-roll the first sentence until it lands; leave a clean close alone.
- 07
Download WAV (or high-bitrate MP3), then drop it on the demo or landing-page MP4
After a generation, ElevenLabs’ help says you can download immediately from the control on the bottom right. Older takes live in History on the Text to Speech page — History lists MP3 (128 kbps) or WAV, with Advanced formats of MP3 192 / 256 kbps, M4A, and FLAC. Higher-quality options are listed on paid tiers; confirm the live download list. Text to Speech is an audio export. We do not invent an MP4 button on that playground. Name the file so the next offer change is obvious:
ledgerline_hero_en_v1. Then put the audio on the video you already have: import the WAV onto a dedicated VO track in your editor, line the first word to the first picture change, leave a breath before a CTA card, and duck any music bed under the narration. Export that timeline as the landing-page or demo MP4. When the price or the hero line changes, edit the sentence, generate, bump to v2, replace the clip. That recut is why you did not book a booth for a thirty-second pitch.
Sales VO vs talking-head demo vs Murf narrator
“AI sales voiceover” is a search, not a product. The decision is the format. The ElevenLabs column matches Text to Speech help and the live pricing page as of August 2026. Synthesia and Murf stay qualitative here — we already wrote those desks, and this page does not hop there.
| Criterion | ElevenLabs sales VO | Synthesia talking-head demo | Murf studio narrator |
|---|---|---|---|
| What the prospect hears / sees | A 15–45s realistic pitch you drop on a landing page or demo timeline | A presenter on camera, then the real product on screen, then a CTA card | A directed studio narrator timed to slides — built for lessons, usable as a read |
| When it is the right buy | The page already has picture. You need a human-sounding close you will recut | The demo needs a face and the UI will ship again this quarter | The picture is a lesson deck and the missing piece is a timed studio read |
| What you re-do when the offer changes | Edit the sentence, Generate Speech, replace the clip on the MP4 timeline | Edit the scene and the screen grab, re-render the talking head | Edit the block, regenerate, replace the audio |
| Tool this desk writes for | ElevenLabs Text to Speech — script, Voice Library, Generate Speech, download | Synthesia — ordinary path: the product-demo how-to. Not a hop on this page | Murf Studio — ordinary path: the training-VO how-to. Not a hop on this page |
| Best 2026 fit | A hero or demo bed on a pricing page that should sound like a person, not an LMS read | A SaaS talking-head demo you will refresh — see the Synthesia product-demo page | A module narrator you time to slides — see the Murf training page |
ElevenLabs
A Text to Speech pitch for a page that already has picture. Free plan to audition a script; a paid plan is the commercial-rights download you can put on a landing page.
The script is a 15–45 second pitch, spoken
Do not write a cold open for a documentary. Do not write a first-second Reel hook and then pad it. Time the copy out loud. Fifteen seconds is enough for the cost of doing nothing and one click. Forty-five seconds is a problem, the product, one proof, and a CTA. Longer than that and you are writing a different video.
A working shape, spoken at a normal pace — roughly half a minute:
Most pricing pages lose the trial in the first sentence. Not because the product is wrong — because nobody said what happens after the click.
This is Ledgerline. One inbox for invoices that already have a due date. You paste the PDF, it files the amount, and you get a reminder the morning it is late — not a dashboard you have to remember to open.
Start the fourteen-day trial. The first invoice you paste is the test.
Spell the words the model should say. “Fourteen-day,” not “14-day.” “PDF,” if you want those letters, or “pee-dee-eff” if you do not. ElevenLabs’ Text to Speech help is blunt about numbers and symbols: write them out, especially on multilingual models, because the same digit is pronounced differently across languages. A product called “Q3 ARR” needs those letters in the script the way you want them heard.
Keep one voice for the hero cut and for the recut. A library narrator that survives your product name is worth more than a cinematic whisper that flubs the trial length. Emotional range is not why you are here. If realism-versus-timeline is the actual question, the comparison is ElevenLabs vs Murf vs Synthesys — ordinary path, no hop. Full product notes live in our ElevenLabs review.
Pick a voice. Generate the pitch. Direct the take.
Open Text to Speech. Select a voice from the control ElevenLabs documents at the bottom left — Default Voices or the Voice Library. Preview. Apply one voice and keep it. A prospect who hears March’s hero and October’s price change should still recognize the person who said the product name.
Clone only if that person has to be you. Instant Voice Cloning is the published fast path from short samples (about 1–2 minutes of clean audio in the cloning help). Professional Voice Cloning is the dedicated model: Creator plan or above, a longer training set (published as 30–180 minutes), and a wait while it fine-tunes. We are not going to invent a clone wizard on this page. If you do not already have the voice in My Voices, use How to clone your voice for YouTube and come back.
Paste the pitch as one short block. Voice, then model, then
settings — ElevenLabs ranks those in that order. Multilingual
v2 is the published stable default and the one we would start
on for a weekly English close. Eleven v3 is the expressive
model: audio tags such as [sighs] or
[clears throat], a 5,000-character cap, and
fewer of the older sliders (Speed, Similarity, and Speaker
Boost are documented as unavailable on v3). Flash models are
the low-latency family; a thirty-second preroll does not need
75ms.
Where the sliders exist, the official starting point is Stability around 50, Similarity around 75, Style exaggeration at 0. Speed defaults to 1.0; the documented range is 0.7 to 1.2. Lower Stability for a livelier take, then generate more than once — the model is nondeterministic. Higher Stability for a straighter closer. Do not invent a slider we cannot see on the public page. If a control is missing on your model, change the sentence and re-roll.
Pacing is the sales-specific job. A dash or em-dash is the documented beat — use it before the product name and before the CTA so the close does not run into the offer. Ellipsis adds hesitation, which a landing-page pitch usually does not want. Nudge Speed if the first sentence eats the proof. On Multilingual v2, Flash v2, and Flash v2.5, ElevenLabs also documents an SSML break tag for a timed pause of up to three seconds. Confirm the live syntax in their Text to Speech help rather than pasting markup we cannot see unchanged. On Eleven v3, use audio tags and punctuation — that model’s help says it does not support SSML break tags.
Press Generate Speech. Listen on headphones. Re-roll the first sentence until it lands. Leave a clean close alone. Two free regenerations of the exact same text and settings are the published allowance; any edit to the copy or the sliders is a new generation.
Download WAV, then put the file on the demo you already have
The download is the gate. ElevenLabs’ Text to Speech help: after you generate, use the download control on the bottom right. Earlier takes sit in History on the same page — sidebar Text to Speech, then the history panel (or the history icon above Generate Speech on a narrow screen). History lists MP3 at 128 kbps or WAV; Advanced adds MP3 192 / 256 kbps, M4A, and FLAC. Higher-quality options are listed on paid tiers. For a landing-page or demo handoff, WAV or a high-bitrate MP3 is the usual file. Confirm the live list on the plan you pay for.
Text to Speech is an audio export. We do not invent an MP4
download on that playground. The MP4 is the page asset you
already have — a product screen grab, a silent demo, a hero
loop — plus this VO. Import the WAV onto a dedicated voice
track. Line the first word to the first picture change. Leave
a breath before a CTA card so the URL is not swallowed. If
you use a music bed, duck it under the narration so the
product name is the loudest thing in the first second. Export
that timeline as the MP4 the landing page or sales email
will play. Name the source audio so the next person can find
it: ledgerline_hero_en_v1. When the trial length
changes, open the sentence, generate, bump to v2, replace the
clip. Leave the proof that is still true. That is the reason
you did not film the founder for the close.
Prospects still deserve a plain-language note when the narration is synthetic. Say it in the page footer or the email. If you also upload the demo as video, use the current platform disclosure for realistic synthetic or altered content. That is separate from the vendor license. You need both.
Commercial rights, the landing page, and the next cut
Read the live page. ElevenLabs meters in character credits across the suite. Free is 10,000 credits a month and no commercial license. Starter is the first paid tier that lists a commercial license and Instant Voice Cloning — $6/month on the public grid we verified in August 2026 (elevenlabs.io/pricing). Creator is the tier that lists Professional Voice Cloning ($22/month on that same grid; annual billing is cheaper). Confirm checkout. We do not print a commission rate.
A 15–45 second pitch is a rounding error on credits. The plan question is the license, not the meter. Sit on a paid plan before the file is the one paid traffic hits. If you later want the founder’s likeness on B-roll demos, that is the clone how-to, not a reason to start this brief in Creator. If the remaining job is a presenter on camera, the Synthesia product-demo page is that brief. If the remaining job is a slide narrator, the Murf training page is that brief. If the remaining job is a fifteen-second show open, the podcast-intro how-to is that brief. The voice-tool comparison lives on ElevenLabs vs Murf vs Synthesys. This page does not hop there. The only Try button here is ElevenLabs.