How to Make an IVR Prompt with ElevenLabs
A practical 2026 path: write a short phone-tree or hold script, pick a Voice Library voice, generate the read in Text to Speech, slow the numbers, download WAV, and drop it in the greeting slot. Built from ElevenLabs' live Text to Speech help and pricing (August 2026).
Most “AI IVR” posts skip the two facts that actually decide the file. First: a twenty-second phone tree is not a YouTube voiceover, and it is not a podcast open that happens to mention a phone. Second: a free-tier preview is not a track you put on a customer-facing line. This guide is the working path for the spoken menu — company name, press one, press two, generate, export WAV, drop it in the greeting slot — using the one voice desk we send for this brief: ElevenLabs.
When a stock ElevenLabs voice is enough — and when you want a booth or a voicemail desk
Use ElevenLabs when the open is a short, clear menu you will regenerate. “For sales, press one.” A Saturday-hours recut. A hold line that still sounds like the person who said the company name. The economic case is the recut: edit the sentence, generate again, replace the file. You do not re-book a booth for twenty seconds.
Skip ElevenLabs if a live receptionist already answers and the words will not change. That file is a recording, then a cut. Skip both if the remaining job is a missed-call drop with a callback number. That is How to make a sales voicemail with Murf — an ordinary guide link, not a second money hop. Murf is the foil, not a Try button. Skip a stock library voice if the tree is a named founder’s voice and you do not have a clone yet. That is How to clone your voice with ElevenLabs — also an ordinary path from this page.
The seven-step ElevenLabs IVR prompt
- 01
Write a 15–25 second phone-tree script — company name, press one / press two, numbers spoken slowly
An IVR prompt is a spoken menu, not a YouTube narration and not a fifteen-second show sting. Time it out loud on a phone speaker in a hallway. Fifteen seconds is the company name plus two options; twenty-five is that plus hours and “press zero to hear this again.” Open on the name the caller already dialed. Give each key its own sentence. Write the digits as words — “one,” “two,” “zero” — so the model does not rush “press 1.” A hold line is a second, shorter file: stay on the line, we will pick up in order. If you wanted a 2–5 minute YouTube episode, that is the YouTube-voiceover how-to. If you wanted a 15–30s Shorts hook, that is the Shorts-voiceover how-to. If you wanted a Day-1 welcome, that is the onboarding how-to. If you wanted a 10–20s show open, that is the podcast-intro how-to.
- 02
Confirm a generated prompt is the right telephony file
Use ElevenLabs when the missing piece is a short, clear menu or hold line you can regenerate when a department name or an hour changes, and the phone tree already exists — a hosted PBX, a ring-group greeting, a hold slot that will not be re-recorded in a booth. Skip it if the remaining job is a missed-call drop with a callback number. That is a voicemail, and the Murf sales-voicemail how-to is an ordinary path on this site, not a hop here. Skip it if a live receptionist already answers and the words will not change. Skip a library voice if callers would take the tree as a named founder. That is a clone, not this brief.
- 03
Open Text to Speech — the speech playground, not Agents
ElevenLabs’ published product for this brief is Text to Speech: paste the menu, pick a voice, generate speech, download a file. That is the desk this page is written for. ElevenAgents is a different product (conversational agents). Studio is a longer project surface for chaptered work. Image & Video on the pricing page is a different surface — we are not inventing an MP4, a phone-tree builder, or an 8 kHz telephony export from the speech playground. This page sends you to ElevenLabs only.
- 04
Sit on a plan that grants commercial rights before the file is the greeting callers hear
ElevenLabs’ own docs are explicit: you keep ownership of generated audio, but commercial usage rights come with paid plans. Free is an audition — published pricing lists 10,000 credits a month and no commercial license. Starter is the first paid tier that lists a commercial license and Instant Voice Cloning. Professional Voice Cloning is listed on Creator and above. Confirm the live grid on elevenlabs.io/pricing. We do not invent a checkout total or a commission rate. Do not publish a free-tier menu on a customer-facing phone tree or a paid support line.
- 05
Pick a Voice Library voice — one calm narrator for the menu and the hold line
Open Voices and browse Default Voices or the Voice Library. Preview before you apply. Cast one narrator and keep it for the greeting and the hold recut. Filter toward clear, professional reads, then listen on a phone speaker — an IVR should sound like a person finishing a sentence in a quiet room, not a cinematic whisper and not a first-second Reel hook. A stock voice is enough when the company has no existing vocal identity. Clone only if callers would notice a stranger. Instant Voice Cloning is the published self-serve path from short samples (ElevenLabs’ cloning help: about 1–2 minutes of clean audio) on paid plans that list it. Professional Voice Cloning trains a dedicated model on a longer set (published as 30–180 minutes) and requires Creator or above. Do not invent a clone wizard here. Do not clone a receptionist from last week’s voicemail dump.
- 06
Paste the menu, pick a model, generate, then slow the options and re-roll
Type or paste the greeting into the text box as one short block. Voice first, then model, then settings — that is the order ElevenLabs’ Text to Speech guide ranks. For a short English phone tree, Multilingual v2 is the published “most stable on long-form” model and still the safest default for a line you will reuse every quarter. Eleven v3 is the expressive model; it supports audio tags such as [sighs] or [clears throat], and it does not expose every older slider — a phone menu usually does not want a sigh. Official starting settings for the sliders that exist: Stability around 50, Similarity around 75, Style exaggeration at 0. Speed defaults to 1.0 (range 0.7–1.2) where the model offers it — nudge Speed down, or add a dash / em-dash, before every “press one.” Spell out numbers. Then press Generate Speech. The model is nondeterministic — same text can yield a different take. Re-roll until each option is unmistakable; leave a clean hold line alone.
- 07
Download WAV, then drop the file in the phone-tree or hold slot you already have
After a generation, ElevenLabs’ help says you can download immediately from the control on the bottom right. Older takes live in History on the Text to Speech page — History lists MP3 (128 kbps) or WAV, with Advanced formats of MP3 192 / 256 kbps, M4A, and FLAC. Higher-quality options are listed on paid tiers; confirm the live download list. Text to Speech is an audio export. We do not invent a PBX preset, an 8 kHz button, or a-LAW / μ-LAW download on that playground. Name the file so the next recut is obvious:
harborline_ivr_sales-support_en_v1. Then put the WAV where the phone system already plays a greeting or a hold prompt. Confirm the live upload rules, sample-rate, and duration cap on the PBX or carrier you use — convert after the download if that system wants a narrower file. When a department name or an hour changes, edit the sentence, generate, bump to v2, replace the clip.
Phone-tree prompt vs a missed-call drop vs a live greeting
“AI IVR prompt” is a search, not a product. The decision is the file. The ElevenLabs column matches Text to Speech help and the live pricing page as of August 2026. A Murf voicemail and a live booth stay qualitative here — we already wrote the outbound desk, and this page does not hop there.
| Criterion | ElevenLabs IVR prompt | Murf sales voicemail | Live receptionist |
|---|---|---|---|
| What the caller hears | A 15–25s phone-tree or hold line you can regenerate when a key or an hour changes | A 20–30s missed-call drop: name, one reason to call back, the number | The actual person, in the actual room, a file that ages when you rename a department |
| When it is the right buy | The tree exists. You need a clear menu you will recut without a booth | The remaining job is an outbound voicemail a dialer will play | Callers would notice a stranger, and the words will not change this quarter |
| What you re-do when the menu changes | Edit the sentence, Generate Speech, replace the file in the greeting slot | Edit the block, regenerate, replace the file in the dialer | Re-book the receptionist, the room, and an editor |
| Tool this desk writes for | ElevenLabs Text to Speech — script, Voice Library, Generate Speech, download WAV | Ordinary path: the Murf sales-voicemail how-to. Not a hop on this page | A booth or a quiet room. Right when the voice has to be live |
| Best 2026 fit | A short inbound menu or hold prompt that should sound like a person, not a stock sting pack | A 20–30s outbound drop — see the Murf sales-voicemail page | A one-time live greeting the front desk already records |
ElevenLabs
A Text to Speech menu for a phone tree that already has a greeting slot. Free plan to audition the options; a paid plan is the commercial-rights download you can put on the line.
The script is a 15–25 second phone tree, spoken slowly
Do not write a cold open for a documentary. Do not write a first-second Reel hook. Do not write a show sting and then tack “press one” on the end. Time the copy out loud. Fifteen seconds is enough for the company name and two keys. Twenty- five seconds is hours and a repeat. Longer than that and you are writing a hold essay the caller will skip.
A working inbound shape, spoken at a slower-than-normal pace — roughly twenty seconds:
Thank you for calling Harborline. Our phones are open Monday through Friday, eight a.m. to six p.m. Eastern.
For sales — press one.
For support — press two.
To hear these options again — press zero.
That is an IVR prompt. The first sentence is the company name. Each option is its own line. The last sentence is a way back to the menu. A trial URL on the end is a landing- page move — if you need a product close, that is a sales voiceover, not this file.
Write the hold line in the same sitting so both files share a voice:
Please stay on the line. A Harborline specialist will be with you shortly. We answer in the order you called.
Spell the words the model should say. “One,” not “1.” “Eight a.m.,” not “8am.” “Eastern,” if that is the timezone you want heard. ElevenLabs’ Text to Speech help is blunt about numbers and symbols: write them out, especially on multilingual models, because the same digit is pronounced differently across languages. A menu about “Q3 ARR” needs those letters in the script the way you want them heard — and if your phone tree is not about that, do not paste a product acronym into a greeting.
Keep one voice for the menu and the hold recut. A library narrator that survives your company name is worth more than a cinematic whisper that swallows “press two.” Emotional range is not why you are here. If realism-versus-timeline is the actual question, the comparison is ElevenLabs vs Murf vs Synthesys — ordinary path, no hop. Full product notes live in our ElevenLabs review.
Pick a voice. Generate the menu. Slow the keys.
Open Text to Speech. Select a voice from the control ElevenLabs documents at the bottom left — Default Voices or the Voice Library. Preview. Apply one voice and keep it. A caller who hears March’s greeting and October’s hours recut should still recognize the person who said the company name.
Clone only if that person has to be a named founder or receptionist. Instant Voice Cloning is the published fast path from short samples (about 1–2 minutes of clean audio in the cloning help). Professional Voice Cloning is the dedicated model: Creator plan or above, a longer training set (published as 30–180 minutes), and a wait while it fine-tunes. We are not going to invent a clone wizard on this page. If callers would notice a stranger and you do not already have the voice in My Voices, use How to clone your voice with ElevenLabs and come back.
Paste the menu as one short block. Voice, then model, then
settings — ElevenLabs ranks those in that order.
Multilingual v2 is the published stable default and the one
we would start on for a quarterly English greeting. Eleven
v3 is the expressive model: audio tags such as
[sighs] or [clears throat], a
5,000-character cap, and fewer of the older sliders (Speed,
Similarity, and Speaker Boost are documented as unavailable
on v3). A phone menu does not need a sigh. Flash models
are the low-latency family; a twenty-second greeting does
not need 75ms.
Where the sliders exist, the official starting point is Stability around 50, Similarity around 75, Style exaggeration at 0. Speed defaults to 1.0; the documented range is 0.7 to 1.2. Lower Stability for a livelier take, then generate more than once — the model is nondeterministic. Higher Stability for a straighter announcer. For a phone tree, start slower than a YouTube open: nudge Speed down until “one” and “two” are unmistakable on a cheap handset. Do not invent a slider we cannot see on the public page. If a control is missing on your model, change the sentence and re-roll.
Pacing is the IVR-specific job. A dash or em-dash is the documented beat — use it before every key so “press one” does not run into “press two.” Ellipsis adds hesitation, which a menu usually does not want. On Multilingual v2, Flash v2, and Flash v2.5, ElevenLabs also documents an SSML break tag for a timed pause of up to three seconds. Confirm the live syntax in their Text to Speech help rather than pasting markup we cannot see unchanged. On Eleven v3, use audio tags and punctuation — that model’s help says it does not support SSML break tags.
Press Generate Speech. Listen on a phone speaker, then on headphones. Re-roll until each option is clear. Leave a clean hold line alone. Two free regenerations of the exact same text and settings are the published allowance; any edit to the copy or the sliders is a new generation.
Download WAV, then put the file where the greeting already lives
The download is the gate. ElevenLabs’ Text to Speech help: after you generate, use the download control on the bottom right. Earlier takes sit in History on the same page — sidebar Text to Speech, then the history panel (or the history icon above Generate Speech on a narrow screen). History lists MP3 at 128 kbps or WAV; Advanced adds MP3 192 / 256 kbps, M4A, and FLAC. Higher-quality options are listed on paid tiers. For a phone-tree or hold-slot handoff, WAV is the usual first file. Confirm the live list on the plan you pay for.
Text to Speech is an audio export. We do not invent a PBX preset, an 8 kHz button, or an a-LAW / μ-LAW download on that playground. The tree is the slot you already have — the hosted greeting, the ring-group prompt, the hold message — plus this file. Import the WAV. Leave a breath after the last key so the carrier does not clip “zero.” If you use a hold bed under the stay-on-the-line line, duck it so the company name is the loudest thing in the first second. We do not invent a carrier preset inside ElevenLabs. Confirm the live upload rules, sample-rate, and duration cap on the PBX or carrier you use, then convert after the download if that system wants a narrower telephony file.
Export is not a second ElevenLabs product. Name the source
audio so the next person can find it:
harborline_ivr_sales-support_en_v1. When a
department name or an hour changes, open the sentence,
generate, bump to v2, replace the file. Leave the hold line
that is still true. That is the reason you did not film the
receptionist for the menu.
Callers still deserve a plain-language note when a person could take the voice as a real staffer speaking. Say it in the internal runbook if not on the line. If you later publish the same file outside the phone system, you also need the current platform disclosure for realistic synthetic or altered content. That is separate from the vendor license. You need both.
Commercial rights, the phone line, and the next cut
Read the live page. ElevenLabs meters in character credits across the suite. Free is 10,000 credits a month and no commercial license. Starter is the first paid tier that lists a commercial license and Instant Voice Cloning — $6/month on the public grid we verified in August 2026 (elevenlabs.io/pricing). Creator is the tier that lists Professional Voice Cloning ($22/month on that same grid; annual billing is cheaper). Confirm checkout. We do not print a commission rate.
A 15–25 second greeting is a rounding error on credits. The plan question is the license, not the meter. Sit on a paid plan before the file is the one callers hear. If you later want a named founder’s likeness on the tree, that is the clone how-to, not a reason to start this brief in Creator. If the remaining job is a four-minute episode, the YouTube-voiceover how-to is that brief. If the remaining job is a 9:16 hook, the Shorts-voiceover how-to is that brief. If the remaining job is a Day-1 welcome, the onboarding-voiceover how-to is that brief. If the remaining job is a fifteen-second show open, the podcast-intro how-to is that brief. If the remaining job is a missed-call drop, the Murf sales-voicemail page is that brief. The voice-tool comparison lives on ElevenLabs vs Murf vs Synthesys. This page does not hop there. The only Try button here is ElevenLabs.
Frequently Asked Questions
What is the best AI tool for an IVR or hold prompt in 2026? +
Can I make an IVR prompt with ElevenLabs for free? +
Should I use a stock ElevenLabs voice or clone the receptionist for the phone tree? +
How do I export an ElevenLabs IVR prompt as WAV for a phone system? +
How do I make “press one” and “press two” easy to hear? +
When should I use ElevenLabs instead of Murf, a live receptionist, or a YouTube-style VO? +
How is this different from the YouTube, Shorts, onboarding, and podcast-intro how-tos? +
Continue the Pipeline
- Guide How to make a YouTube voiceover with ElevenLabs →
- Guide How to make a YouTube Shorts voiceover with ElevenLabs →
- Guide How to make an onboarding voiceover with ElevenLabs →
- Guide How to make a podcast intro with ElevenLabs →
- Guide How to make a sales voicemail with Murf →
- Guide How to clone your voice with ElevenLabs →
- Review ElevenLabs voice cloning review →