How to Make a Podcast Intro with ElevenLabs
A practical 2026 path: write a 10–20 second show-open line, pick a Voice Library voice, generate the sting in Text to Speech, download WAV, and drop it at the top of the episode. Built from ElevenLabs' live Text to Speech help and pricing (August 2026).
Most “AI podcast intro” posts skip the two facts that actually decide the file. First: a fifteen-second show open is not a sales close, and it is not a music sting that tries to say the title for you. Second: a free-tier preview is not a track you put on a monetized feed. This guide is the working path for the spoken open — one line, one voice, generate, export, drop it on the episode — using the one voice desk we send for this brief: ElevenLabs.
When a stock ElevenLabs voice is enough — and when you want a booth or a bed
Use ElevenLabs when the open is a short, realistic line you will regenerate. “You’re listening to Hold for Sound.” A new season title. A closer that still sounds like the person who said the show name. The economic case is the recut: edit the sentence, generate again, replace the file. You do not re-book a booth for fifteen seconds.
Skip ElevenLabs if the host already records a live open every week and the words will not change. That file is a recording, then a cut. Skip both if the spoken open already exists and the remaining job is a soundtrack. That is How to add AI music to a podcast trailer — an ordinary guide link, not a second money hop. Mubert is the foil, not a Try button. Skip a stock library voice if the show is a named host’s voice and you do not have a clone yet. That is How to clone your voice with ElevenLabs — also an ordinary path from this page.
The seven-step ElevenLabs podcast intro
- 01
Write one 10–20 second show-open line — title, listener, one promise
A podcast intro is a spoken sting, not a landing-page close and not a first-second Short. Time it out loud: ten seconds is the show name plus who it is for; twenty is that plus one sentence of what this episode (or this season) is for. Open on the title the way it should be heard. Name the listener. Stop before you start the interview. Spell numbers, season titles, and brand names in the script. Write a matching closer in the same sitting — thank you, where to subscribe, next week — so both files share one voice. If you wanted a 15–45s product pitch, that is the sales-voiceover how-to. If you wanted a 15–30s Shorts hook, that is the Shorts-voiceover how-to.
- 02
Confirm a generated sting is the right file
Use ElevenLabs when the missing piece is a short, realistic open you can regenerate when the season title or the guest line changes. Skip it if the host already records a live open in the booth and the words will not change — record that, then cut it. Skip it if the remaining job is a music bed under a listen-ask. That is a soundtrack, and the podcast-trailer music how-to is an ordinary path on this site, not a hop here. Skip a library voice if the show is a named host and subscribers would notice a stranger. That is a clone, not this brief.
- 03
Open Text to Speech — the speech playground, not Agents
ElevenLabs’ published product for this brief is Text to Speech: paste the open, pick a voice, generate speech, download a file. That is the desk this page is written for. ElevenAgents is a different product (conversational agents). Studio is a longer project surface for chaptered work. Image & Video on the pricing page is a different surface — we are not inventing an MP4 download from the speech playground. This page sends you to ElevenLabs only.
- 04
Sit on a plan that grants commercial rights before the sting is the feed
ElevenLabs’ own docs are explicit: you keep ownership of generated audio, but commercial usage rights come with paid plans. Free is an audition — published pricing lists 10,000 credits a month and no commercial license. Starter is the first paid tier that lists a commercial license and Instant Voice Cloning. Professional Voice Cloning is listed on Creator and above. Confirm the live grid on elevenlabs.io/pricing. We do not invent a checkout total or a commission rate. Do not publish a free-tier open on a monetized RSS feed or a YouTube episode upload.
- 05
Pick a Voice Library voice — one narrator for the open and the closer
Open Voices and browse Default Voices or the Voice Library. Preview before you apply. Cast one narrator and keep it for the intro and the outro. Filter toward narration / podcast-style reads, then listen on headphones and in a car — that is how most subscribers hear the first second. A stock voice is enough when the show has no existing vocal identity. Clone only if listeners would notice a stranger. Instant Voice Cloning is the published self-serve path from short samples (ElevenLabs’ cloning help: about 1–2 minutes of clean audio) on paid plans that list it. Professional Voice Cloning trains a dedicated model on a longer set (published as 30–180 minutes) and requires Creator or above. Do not invent a clone wizard here. Do not clone a guest from last week’s Riverside or Zoom file.
- 06
Paste the open, pick a model, generate, then re-roll the first three words
Type or paste the sting into the text box as one short block. Voice first, then model, then settings — that is the order ElevenLabs’ Text to Speech guide ranks. For a short English show open, Multilingual v2 is the published “most stable on long-form” model and still the safest default for a line you will reuse every week. Eleven v3 is the expressive model; it supports audio tags such as [sighs] or [clears throat], and it does not expose every older slider. Official starting settings for the sliders that exist: Stability around 50, Similarity around 75, Style exaggeration at 0. Speed defaults to 1.0 (range 0.7–1.2) where the model offers it — nudge Speed, or add a dash / em-dash, when the title feels swallowed. Spell out numbers. Then press Generate Speech. The model is nondeterministic — same text can yield a different take. Re-roll the first three words until they punch; leave a clean closer alone.
- 07
Download WAV or high-bitrate MP3, then drop the file at the top of the episode
After a generation, ElevenLabs’ help says you can download immediately from the control on the bottom right. Older takes live in History on the Text to Speech page — History lists MP3 (128 kbps) or WAV, with Advanced formats of MP3 192 / 256 kbps, M4A, and FLAC. Higher-quality options are listed on paid tiers; confirm the live download list. Text to Speech is an audio export. We do not invent an MP4 button on that playground. Name the file so the next season recut is obvious:
hold-for-sound_intro_en_v1. Then put it on the episode you already have: import the sting into the same project as the interview and lay it before the conversation. In any other editor, drop the WAV or MP3 on a dedicated intro track, leave a breath before the host speaks, and duck any music bed under the VO. When the season name changes, edit the sentence, generate, bump to v2, replace the file. That recut is why you did not book a booth for fifteen seconds.
Spoken sting vs a live booth open vs a music bed
“AI podcast intro” is a search, not a product. The decision is the file. The ElevenLabs column matches Text to Speech help and the live pricing page as of August 2026. A live booth and a Mubert bed stay qualitative here — we already wrote the music desk, and this page does not hop there.
| Criterion | ElevenLabs spoken sting | Record the host | Mubert music bed |
|---|---|---|---|
| What the subscriber hears | A 10–20s spoken open you can regenerate when the season title changes | The actual host, in the actual room, a file that ages when you rename the show | A licensed bed under a listen-ask — music, not a spoken title |
| When it is the right buy | You need a human-sounding show name and you will recut it | The open already happens live and the words will not be rewritten | The spoken open exists and the remaining job is a soundtrack |
| What you re-do when the copy changes | Edit the sentence, Generate Speech, replace the file on the timeline | Re-book the host, the room, and an editor | Generate a new loop, duck it under the VO, keep the certificate |
| Tool this desk writes for | ElevenLabs Text to Speech — script, Voice Library, Generate Speech, download | A booth or a quiet room. Right when the voice has to be live | Ordinary path: the podcast-trailer music how-to. Not a hop on this page |
| Best 2026 fit | A weekly show open/close that should sound like a person, not a stock sting pack | A one-time live open the host already records every week | A 30–60s listen-ask that needs a floor under the host — see the Mubert trailer page |
ElevenLabs
A Text to Speech sting for a show that already has an episode. Free plan to audition the open; a paid plan is the commercial-rights download you can put on the feed.
The script is one show-open line, spoken
Do not write a cold open for a documentary. Do not write a first-second Reel hook. Do not write a product close for a pricing page. Time the copy out loud. Ten seconds is enough for the title and who it is for. Twenty seconds is one promise. Longer than that and you are writing the episode.
A working shape, spoken at a normal pace — roughly a quarter minute:
You’re listening to Hold for Sound — the weekly show for people who still ride a mix after the cut is locked. I’m Priya Shah. This week: one messy room, one decision that stayed.
That is a podcast intro. The first sentence is the show name. The middle names the listener. The last sentence is one promise for this episode. A trial URL on the end is a landing-page move — if you need a sales close, that is the sales-voiceover how-to, not this file.
Write the closer in the same sitting so both files share a voice:
That’s Hold for Sound. If this saved you a pass, send it to the person still in the session. See you next week.
Spell the words the model should say. “Season three,” not “S3.” “Friday,” not “Fri.” ElevenLabs’ Text to Speech help is blunt about numbers and symbols: write them out, especially on multilingual models, because the same digit is pronounced differently across languages. A show called “Q3 ARR” needs those letters in the script the way you want them heard.
Keep one voice for the intro and the outro. A library narrator that survives your show name is worth more than a cinematic whisper that flubs the title. Emotional range is not why you are here. If realism-versus-timeline is the actual question, the comparison is ElevenLabs vs Murf vs Synthesys — ordinary path, no hop. Full product notes live in our ElevenLabs review.
Pick a voice. Generate the open. Export the take.
Open Text to Speech. Select a voice from the control ElevenLabs documents at the bottom left — Default Voices or the Voice Library. Preview. Apply one voice and keep it. A subscriber who hears March’s open and October’s recut should still recognize the person who said the show name.
Clone only if that person has to be you. Instant Voice Cloning is the published fast path from short samples (about 1–2 minutes of clean audio in the cloning help). Professional Voice Cloning is the dedicated model: Creator plan or above, a longer training set (published as 30–180 minutes), and a wait while it fine-tunes. We are not going to invent a clone wizard on this page. If listeners would notice a stranger and you do not already have the voice in My Voices, use How to clone your voice with ElevenLabs and come back.
Paste the open as one short block. Voice, then model, then
settings — ElevenLabs ranks those in that order. Multilingual
v2 is the published stable default and the one we would start
on for a weekly English sting. Eleven v3 is the expressive
model: audio tags such as [sighs] or
[clears throat], a 5,000-character cap, and
fewer of the older sliders (Speed, Similarity, and Speaker
Boost are documented as unavailable on v3). Flash models are
the low-latency family; a fifteen-second preroll does not
need 75ms.
Where the sliders exist, the official starting point is Stability around 50, Similarity around 75, Style exaggeration at 0. Speed defaults to 1.0; the documented range is 0.7 to 1.2. Lower Stability for a livelier take, then generate more than once — the model is nondeterministic. Higher Stability for a straighter announcer. Do not invent a slider we cannot see on the public page. If a control is missing on your model, change the sentence and re-roll.
Pacing is the sting-specific job. A dash or em-dash is the documented beat — use it after the title and before the promise so the open does not run into the guest name. Ellipsis adds hesitation, which a show open usually does not want. Nudge Speed if the first three words eat the title. On Multilingual v2, Flash v2, and Flash v2.5, ElevenLabs also documents an SSML break tag for a timed pause of up to three seconds. Confirm the live syntax in their Text to Speech help rather than pasting markup we cannot see unchanged. On Eleven v3, use audio tags and punctuation — that model’s help says it does not support SSML break tags.
Press Generate Speech. Listen on headphones, then in a car or on a phone speaker. Re-roll the first three words until they punch. Leave a clean closer alone. Two free regenerations of the exact same text and settings are the published allowance; any edit to the copy or the sliders is a new generation.
Download the file, then put it where the episode already lives
The download is the gate. ElevenLabs’ Text to Speech help: after you generate, use the download control on the bottom right. Earlier takes sit in History on the same page — sidebar Text to Speech, then the history panel (or the history icon above Generate Speech on a narrow screen). History lists MP3 at 128 kbps or WAV; Advanced adds MP3 192 / 256 kbps, M4A, and FLAC. Higher-quality options are listed on paid tiers. For a podcast enclosure, WAV or a high-bitrate MP3 is the usual handoff. Confirm the live list on the plan you pay for.
Text to Speech is an audio export. We do not invent an MP4 download on that playground. The episode is the file you already have — the interview, the remote, the tape — plus this sting. Import the WAV onto a dedicated intro track. Line the first word to the first picture change if you also upload video. Leave a breath before the host speaks so the title is not swallowed. If you use a music bed, duck it under the VO so the show name is the loudest thing in the first second. If you still need that bed written, the working path is How to add AI music to a podcast trailer — ordinary page, no hop from here.
Export is not a second ElevenLabs product. Name the source
audio so the next person can find it:
hold-for-sound_intro_en_v1. When the season
title changes, open the sentence, generate, bump to v2,
replace the file. Leave the closer that is still true. That
is the reason you did not film the host for the open.
Listeners still deserve a plain-language note when the open is synthetic. Say it in the show notes. If you also upload the episode as video, use the current platform disclosure for realistic synthetic or altered content. That is separate from the vendor license. You need both.
Commercial rights, the feed, and the next cut
Read the live page. ElevenLabs meters in character credits across the suite. Free is 10,000 credits a month and no commercial license. Starter is the first paid tier that lists a commercial license and Instant Voice Cloning — $6/month on the public grid we verified in August 2026 (elevenlabs.io/pricing). Creator is the tier that lists Professional Voice Cloning ($22/month on that same grid; annual billing is cheaper). Confirm checkout. We do not print a commission rate.
A weekly 15-second sting is a rounding error on credits. The plan question is the license, not the meter. Sit on a paid plan before the file is the one subscribers get. If you later want the host’s likeness on B-roll episodes, that is the clone how-to, not a reason to start this brief in Creator. If the remaining job is a thirty-second product close, the sales-voiceover how-to is that brief. If the remaining job is a company-page post, the LinkedIn-voiceover how-to is that brief. If the remaining job is a Day-1 welcome, the onboarding-voiceover how-to is that brief. If the remaining job is a 9:16 hook, the Shorts-voiceover how-to is that brief. If the remaining job is a four-minute episode, the YouTube-voiceover how-to is that brief. If the remaining job is a listen-ask soundtrack, the Mubert podcast-trailer page is that brief. The voice-tool comparison lives on ElevenLabs vs Murf vs Synthesys. This page does not hop there. The only Try button here is ElevenLabs.
Frequently Asked Questions
What is the best AI tool for a podcast intro in 2026? +
Can I make a podcast intro with ElevenLabs for free? +
Should I use a stock ElevenLabs voice or clone the host for the intro? +
How do I export an ElevenLabs podcast intro and drop it on the episode? +
How do I tweak pacing on a 10–20 second show open? +
How is this different from the sales, LinkedIn, onboarding, Shorts, and YouTube voiceover how-tos? +
Do I need to disclose that the podcast intro is AI-generated? +
Continue the Pipeline
- Guide How to make a sales voiceover with ElevenLabs →
- Guide How to make a LinkedIn voiceover with ElevenLabs →
- Guide How to make an onboarding voiceover with ElevenLabs →
- Guide How to make a YouTube Shorts voiceover with ElevenLabs →
- Guide How to make a YouTube voiceover with ElevenLabs →
- Guide How to add AI music to a podcast trailer →
- Guide How to clone your voice with ElevenLabs →
- Review ElevenLabs voice cloning review →