Skip to content
AI Video Tools Guide
Desk /
Menu
Guide · Audio & Voice Verified August 2026

How to Make a Sales Voiceover with ElevenLabs

A practical 2026 path: write a 15–45 second sales or product-demo script, pick a Voice Library voice, generate the read in Text to Speech, tweak pacing, download WAV, and drop it on the demo or landing-page timeline. Built from ElevenLabs' live Text to Speech help and pricing (August 2026).

By Scott /11 min read

Most “AI sales voiceover” posts skip the two facts that actually decide the file. First: a thirty-second pitch is not a podcast sting, and it is not a talking-head demo. Second: a free-tier preview is not a track you put on a paid-media landing page. This guide is the working path for the spoken close — script, voice, generate, pace, download, drop it on the demo — using the one voice desk we send for this brief: ElevenLabs.

When a stock ElevenLabs voice is enough — and when you want a face or a booth

Use ElevenLabs when the open is a short, realistic read you will regenerate. A hero line on a pricing page. A product-demo bed under a screen grab. A sales-email clip that still sounds like the same person who closed last month’s offer. The economic case is the recut: edit the sentence, generate again, replace the clip. You do not re-book a booth for thirty seconds.

Skip ElevenLabs if the landing page needs a presenter walking the product on camera. That file is a talking-head demo, and the working path on this site is How to make product demo videos with Synthesia — an ordinary guide link, not a second money hop. Skip both if the missing piece is a directed studio narrator on slides. That is How to make training voiceovers with Murf. Murf is the foil, not a Try button. Skip a stock library voice if the brand is a named founder’s voice and you do not have a clone yet. That is How to clone your voice for YouTube.

The seven-step ElevenLabs sales voiceover

  1. 01

    Write a 15–45 second sales script — problem, product, one proof, one CTA

    A sales voiceover is a spoken pitch, not a show sting and not a two-minute explainer. Time it out loud: fifteen seconds is a hero-line plus a click; forty-five is a problem, the product name, one proof, and one next step. Open on the buyer’s cost of doing nothing, name the product, give one number or one outcome a prospect can check, then say the URL or the trial. Spell plan names, dollar figures, and abbreviations the way they should be heard on a landing page. If you wanted a 10–20s show open, that is the podcast-intro how-to. If you wanted a 2–5 minute YouTube narration, that is the YouTube-voiceover how-to.

  2. 02

    Confirm a generated VO is the right sales format

    Use ElevenLabs when the missing piece is a short, realistic read you can regenerate when the offer or the hero line changes. Skip it if the landing page needs a presenter on camera walking the product. That is a talking-head demo, and that how-to is an ordinary path on this site, not a hop here. Skip it if the file is a directed studio narrator timed to a slide lesson. That is Murf Studio, and that how-to is also an ordinary path. Skip a library voice if the brand is the founder’s voice and buyers would notice a stranger. That is a clone, and the full training-set path lives on the YouTube clone how-to.

  3. 03

    Open Text to Speech — the speech playground, not Agents

    ElevenLabs’ published product for this brief is Text to Speech: paste text, pick a voice, generate speech, download a file. That is the desk this page is written for. ElevenAgents is a different product (conversational agents). Studio is a longer project surface for chaptered work. Image & Video on the pricing page is a different surface — we are not inventing an MP4 download from the speech playground. This page sends you to ElevenLabs only.

  4. 04

    Sit on a plan that grants commercial rights before you treat the file as the landing page

    ElevenLabs’ own docs are explicit: you keep ownership of generated audio, but commercial usage rights come with paid plans. Free is an audition — published pricing lists 10,000 credits a month and no commercial license. Starter is the first paid tier that lists a commercial license and Instant Voice Cloning. Professional Voice Cloning is listed on Creator and above. Confirm the live grid on elevenlabs.io/pricing. We do not invent a checkout total or a commission rate. Do not publish a free-tier read on a paid-media landing page or a sales email.

  5. 05

    Pick a Voice Library voice — clone only if the brand already is your voice

    Open Voices and browse Default Voices or the Voice Library. Preview before you apply. Cast one narrator and keep it for the landing-page cut, the recut, and any matching demo bed. Filter toward narration-style reads, then listen on headphones — a sales VO should sound like a calm closer, not a cinematic whisper and not a training-module drone. A stock voice is enough when the brand has no existing vocal identity. Clone only if buyers would notice a stranger. Instant Voice Cloning is the published self-serve path from short samples (ElevenLabs’ cloning help: about 1–2 minutes of clean audio) on paid plans that list it. Professional Voice Cloning trains a dedicated model on a longer set (published as 30–180 minutes) and requires Creator or above. Do not invent a clone wizard here — if you need the training-set workflow, use the clone how-to. Do not clone a customer, a podcast guest, or a voice you found on someone else’s demo.

  6. 06

    Paste the 15–45s script, pick a model, generate, then tweak pacing and re-roll

    Type or paste the pitch into the text box as one short block. Voice first, then model, then settings — that is the order ElevenLabs’ Text to Speech guide ranks. For a short English sales read, Multilingual v2 is the published “most stable on long-form” model and still the safest default for a line you will reuse on a landing page. Eleven v3 is the expressive model; it supports audio tags such as [sighs] or [clears throat], and it does not expose every older slider. Official starting settings for the sliders that exist: Stability around 50, Similarity around 75, Style exaggeration at 0. Speed defaults to 1.0 (range 0.7–1.2) where the model offers it — nudge Speed, or add a dash / em-dash, when the close feels rushed. Spell out numbers. Then press Generate Speech. The model is nondeterministic — same text can yield a different take. Re-roll the first sentence until it lands; leave a clean close alone.

  7. 07

    Download WAV (or high-bitrate MP3), then drop it on the demo or landing-page MP4

    After a generation, ElevenLabs’ help says you can download immediately from the control on the bottom right. Older takes live in History on the Text to Speech page — History lists MP3 (128 kbps) or WAV, with Advanced formats of MP3 192 / 256 kbps, M4A, and FLAC. Higher-quality options are listed on paid tiers; confirm the live download list. Text to Speech is an audio export. We do not invent an MP4 button on that playground. Name the file so the next offer change is obvious: ledgerline_hero_en_v1. Then put the audio on the video you already have: import the WAV onto a dedicated VO track in your editor, line the first word to the first picture change, leave a breath before a CTA card, and duck any music bed under the narration. Export that timeline as the landing-page or demo MP4. When the price or the hero line changes, edit the sentence, generate, bump to v2, replace the clip. That recut is why you did not book a booth for a thirty-second pitch.

Sales VO vs talking-head demo vs Murf narrator

“AI sales voiceover” is a search, not a product. The decision is the format. The ElevenLabs column matches Text to Speech help and the live pricing page as of August 2026. Synthesia and Murf stay qualitative here — we already wrote those desks, and this page does not hop there.

Sales audio: ElevenLabs VO vs Synthesia talking-head demo vs Murf narrator (August 2026)
Criterion ElevenLabs sales VOSynthesia talking-head demoMurf studio narrator
What the prospect hears / sees A 15–45s realistic pitch you drop on a landing page or demo timeline A presenter on camera, then the real product on screen, then a CTA card A directed studio narrator timed to slides — built for lessons, usable as a read
When it is the right buy The page already has picture. You need a human-sounding close you will recut The demo needs a face and the UI will ship again this quarter The picture is a lesson deck and the missing piece is a timed studio read
What you re-do when the offer changes Edit the sentence, Generate Speech, replace the clip on the MP4 timeline Edit the scene and the screen grab, re-render the talking head Edit the block, regenerate, replace the audio
Tool this desk writes for ElevenLabs Text to Speech — script, Voice Library, Generate Speech, download Synthesia — ordinary path: the product-demo how-to. Not a hop on this page Murf Studio — ordinary path: the training-VO how-to. Not a hop on this page
Best 2026 fit A hero or demo bed on a pricing page that should sound like a person, not an LMS read A SaaS talking-head demo you will refresh — see the Synthesia product-demo page A module narrator you time to slides — see the Murf training page
The sales-VO desk

ElevenLabs

A Text to Speech pitch for a page that already has picture. Free plan to audition a script; a paid plan is the commercial-rights download you can put on a landing page.

The script is a 15–45 second pitch, spoken

Do not write a cold open for a documentary. Do not write a first-second Reel hook and then pad it. Time the copy out loud. Fifteen seconds is enough for the cost of doing nothing and one click. Forty-five seconds is a problem, the product, one proof, and a CTA. Longer than that and you are writing a different video.

A working shape, spoken at a normal pace — roughly half a minute:

Most pricing pages lose the trial in the first sentence. Not because the product is wrong — because nobody said what happens after the click.

This is Ledgerline. One inbox for invoices that already have a due date. You paste the PDF, it files the amount, and you get a reminder the morning it is late — not a dashboard you have to remember to open.

Start the fourteen-day trial. The first invoice you paste is the test.

Spell the words the model should say. “Fourteen-day,” not “14-day.” “PDF,” if you want those letters, or “pee-dee-eff” if you do not. ElevenLabs’ Text to Speech help is blunt about numbers and symbols: write them out, especially on multilingual models, because the same digit is pronounced differently across languages. A product called “Q3 ARR” needs those letters in the script the way you want them heard.

Keep one voice for the hero cut and for the recut. A library narrator that survives your product name is worth more than a cinematic whisper that flubs the trial length. Emotional range is not why you are here. If realism-versus-timeline is the actual question, the comparison is ElevenLabs vs Murf vs Synthesys — ordinary path, no hop. Full product notes live in our ElevenLabs review.

Pick a voice. Generate the pitch. Direct the take.

Open Text to Speech. Select a voice from the control ElevenLabs documents at the bottom left — Default Voices or the Voice Library. Preview. Apply one voice and keep it. A prospect who hears March’s hero and October’s price change should still recognize the person who said the product name.

Clone only if that person has to be you. Instant Voice Cloning is the published fast path from short samples (about 1–2 minutes of clean audio in the cloning help). Professional Voice Cloning is the dedicated model: Creator plan or above, a longer training set (published as 30–180 minutes), and a wait while it fine-tunes. We are not going to invent a clone wizard on this page. If you do not already have the voice in My Voices, use How to clone your voice for YouTube and come back.

Paste the pitch as one short block. Voice, then model, then settings — ElevenLabs ranks those in that order. Multilingual v2 is the published stable default and the one we would start on for a weekly English close. Eleven v3 is the expressive model: audio tags such as [sighs] or [clears throat], a 5,000-character cap, and fewer of the older sliders (Speed, Similarity, and Speaker Boost are documented as unavailable on v3). Flash models are the low-latency family; a thirty-second preroll does not need 75ms.

Where the sliders exist, the official starting point is Stability around 50, Similarity around 75, Style exaggeration at 0. Speed defaults to 1.0; the documented range is 0.7 to 1.2. Lower Stability for a livelier take, then generate more than once — the model is nondeterministic. Higher Stability for a straighter closer. Do not invent a slider we cannot see on the public page. If a control is missing on your model, change the sentence and re-roll.

Pacing is the sales-specific job. A dash or em-dash is the documented beat — use it before the product name and before the CTA so the close does not run into the offer. Ellipsis adds hesitation, which a landing-page pitch usually does not want. Nudge Speed if the first sentence eats the proof. On Multilingual v2, Flash v2, and Flash v2.5, ElevenLabs also documents an SSML break tag for a timed pause of up to three seconds. Confirm the live syntax in their Text to Speech help rather than pasting markup we cannot see unchanged. On Eleven v3, use audio tags and punctuation — that model’s help says it does not support SSML break tags.

Press Generate Speech. Listen on headphones. Re-roll the first sentence until it lands. Leave a clean close alone. Two free regenerations of the exact same text and settings are the published allowance; any edit to the copy or the sliders is a new generation.

Download WAV, then put the file on the demo you already have

The download is the gate. ElevenLabs’ Text to Speech help: after you generate, use the download control on the bottom right. Earlier takes sit in History on the same page — sidebar Text to Speech, then the history panel (or the history icon above Generate Speech on a narrow screen). History lists MP3 at 128 kbps or WAV; Advanced adds MP3 192 / 256 kbps, M4A, and FLAC. Higher-quality options are listed on paid tiers. For a landing-page or demo handoff, WAV or a high-bitrate MP3 is the usual file. Confirm the live list on the plan you pay for.

Text to Speech is an audio export. We do not invent an MP4 download on that playground. The MP4 is the page asset you already have — a product screen grab, a silent demo, a hero loop — plus this VO. Import the WAV onto a dedicated voice track. Line the first word to the first picture change. Leave a breath before a CTA card so the URL is not swallowed. If you use a music bed, duck it under the narration so the product name is the loudest thing in the first second. Export that timeline as the MP4 the landing page or sales email will play. Name the source audio so the next person can find it: ledgerline_hero_en_v1. When the trial length changes, open the sentence, generate, bump to v2, replace the clip. Leave the proof that is still true. That is the reason you did not film the founder for the close.

Prospects still deserve a plain-language note when the narration is synthetic. Say it in the page footer or the email. If you also upload the demo as video, use the current platform disclosure for realistic synthetic or altered content. That is separate from the vendor license. You need both.

Commercial rights, the landing page, and the next cut

Read the live page. ElevenLabs meters in character credits across the suite. Free is 10,000 credits a month and no commercial license. Starter is the first paid tier that lists a commercial license and Instant Voice Cloning — $6/month on the public grid we verified in August 2026 (elevenlabs.io/pricing). Creator is the tier that lists Professional Voice Cloning ($22/month on that same grid; annual billing is cheaper). Confirm checkout. We do not print a commission rate.

A 15–45 second pitch is a rounding error on credits. The plan question is the license, not the meter. Sit on a paid plan before the file is the one paid traffic hits. If you later want the founder’s likeness on B-roll demos, that is the clone how-to, not a reason to start this brief in Creator. If the remaining job is a presenter on camera, the Synthesia product-demo page is that brief. If the remaining job is a slide narrator, the Murf training page is that brief. If the remaining job is a fifteen-second show open, the podcast-intro how-to is that brief. The voice-tool comparison lives on ElevenLabs vs Murf vs Synthesys. This page does not hop there. The only Try button here is ElevenLabs.

Frequently Asked Questions

What is the best AI tool for a sales or product-demo voiceover in 2026? +
ElevenLabs Text to Speech, when the job is a 15–45 second sales or product-demo read that has to sound human and you will regenerate it when the offer changes. Write the pitch, pick a Voice Library voice (or an existing clone if the brand is already your voice), press Generate Speech, download WAV or a high-bitrate MP3, and drop the file on the demo or landing-page timeline — then export that timeline as the MP4 the page plays. Synthesia is the better pick when the missing piece is a presenter on camera walking the product — that is the product-demo how-to, an ordinary path on this site. Murf is the better pick when the picture is a lesson deck — that is the training-VO how-to, also an ordinary path. Record the founder if the close already happens live and will not be rewritten. Buying a talking-head suite for audio-only, or booking a half-day to rerecord “fourteen-day trial,” is the expensive mistake.
Can I make a sales voiceover with ElevenLabs for free? +
You can audition. ElevenLabs’ published Free plan is 10,000 credits a month. That is enough to hear whether a library voice survives your product name and a real CTA. It is not a commercial license. ElevenLabs’ own docs: you retain ownership of generated audio, but commercial usage rights are only available with paid plans. Sit on a paid plan before the file is the one on a paid-media landing page or in a sales email. Confirm the live grid on elevenlabs.io/pricing. We do not invent a checkout total.
Should I use a stock ElevenLabs voice or clone the founder for a sales VO? +
Pick a library voice unless buyers would notice a stranger. Most landing-page pitches have no existing vocal identity: one consistent narrator for the hero cut and for later offer recuts is the product. Clone only if the brand already is the founder’s voice. Instant Voice Cloning is the published self-serve path from about 1–2 minutes of clean audio on plans that list it. Professional Voice Cloning trains a dedicated model on a longer set (published as 30–180 minutes) and requires Creator or above. The full consent-and-training workflow lives at How to clone your voice for YouTube. Do not clone a customer testimonial or a voice you found on someone else’s demo.
How do I export a sales voiceover as WAV or MP4 from ElevenLabs? +
Text to Speech downloads audio, not a finished landing-page video. After Generate Speech, use the download control on the bottom right, or pull the take from History as MP3 (128 kbps) or WAV. Advanced formats listed in ElevenLabs’ help are MP3 192 / 256 kbps, M4A, and FLAC — higher-quality options sit on paid tiers; confirm the live list. For a demo or pricing-page handoff, WAV or a high-bitrate MP3 is the usual file. Import that file onto a dedicated VO track in your editor, line it to picture, and export the timeline as the MP4 the page or email will play. We do not invent an MP4 button on the speech playground. If the brief is a talking-head MP4 with a presenter on camera, that is How to make product demo videos with Synthesia — ordinary path, no hop from here.
How do I tweak pacing on a 15–45 second sales read? +
Start with the official sliders that exist on your model: Speed defaults to 1.0 and is documented from 0.7 to 1.2 where the model offers it. Nudge Speed down if the CTA feels swallowed; nudge it up if the problem line drags. Punctuation is the other pacing tool — a dash or em-dash is the documented beat; ellipsis adds hesitation, which a close usually does not want. On Multilingual v2, Flash v2, and Flash v2.5, ElevenLabs also documents an SSML break tag for a timed pause of up to three seconds. Confirm the live syntax in their Text to Speech help rather than pasting markup we cannot see unchanged. On Eleven v3, use audio tags and punctuation — that model’s help says it does not support SSML break tags. If a control is missing on your model, change the sentence and re-roll. Two free regenerations of the exact same text and settings are the published allowance; any edit to the copy or the sliders is a new generation.
When should I use ElevenLabs instead of Synthesia or Murf — or instead of recording the founder? +
Use ElevenLabs when the deliverable is a short, realistic sales read you will recut without a booth or a face. Use Synthesia when the demo needs a presenter on camera plus a screen grab of the real product. That is How to make product demo videos with Synthesia, an ordinary site path, not a second money button. Use Murf when the picture already exists on slides and you need a directed studio narrator on a timeline. That is How to make training voiceovers with Murf, also an ordinary path. Record the founder when the close is already a live performance and the words will not change. Same launch calendar. Different file. Full voice-tool scorecard: ElevenLabs vs Murf vs Synthesys.
How is this different from the podcast-intro how-to and the YouTube-voiceover how-to? +
The podcast-intro page is a 10–20 second show open: show name, who it is for, one promise, drop it on the episode. The YouTube-voiceover page is a 2–5 minute spoken episode: hook, two or three beats, close, generate in chunks. This page is the 15–45 second sales or product-demo pitch: problem, product, one proof, one CTA, download the audio, drop it on a landing-page or demo MP4. Same company on the hop as other ElevenLabs pages. Different brief. If the file is a fifteen-second show sting, start at the podcast-intro how-to. If the file is a four-minute explainer, start at the YouTube-voiceover how-to. If the file should look like a presenter walking the product, start at the Synthesia product-demo how-to.

Continue the Pipeline

Sponsored

Try ElevenLabs