Skip to content
AI Video Tools Guide
Desk /
Menu
Guide · Audio & Voice Verified August 2026

How to Make a Podcast Intro with ElevenLabs

A practical 2026 path: write a 10–20 second intro (and a matching outro), pick a voice or reuse a clone, generate the read in Text to Speech, download the file, and drop it at the top of the episode in Descript or your editor. Built from ElevenLabs' live Text to Speech help and pricing (August 2026).

By Scott /11 min read

Most “AI podcast intro” posts skip the two facts that actually decide the file. First: a ten-second show open is not a cloned host, and it is not the radio edit of the interview. Second: a free-tier preview is not a sting you put on a monetized feed. This guide is the working path for the spoken open — script, voice, generate, download, drop it on the episode — using the one voice desk we send for this brief: ElevenLabs.

When ElevenLabs is the VO tool — and when you want Murf or a booth

Use ElevenLabs when the open is a short, realistic read you will regenerate. “You’re listening to After the Cut.” A new season title. A sponsor-safe closer that still sounds like the same person who opened the show. The economic case is the recut: edit the sentence, generate again, replace the file. You do not re-book a booth for fifteen seconds.

Skip ElevenLabs if the host already records a live open every week and the words will not change. That file is a recording, then a cut. The working path on this site for the episode around it is How to edit a podcast in Descript — an ordinary guide link, not a second money hop. Skip both if the missing piece is a directed studio narrator on slides. That is How to make training voiceovers with Murf. Murf is the foil, not a Try button. Skip a stock library voice if the show is a named host’s voice and you do not have a clone yet. That is How to clone your voice for YouTube.

The seven-step ElevenLabs podcast intro

  1. 01

    Write a 10–20 second intro — show name, who it is for, one promise

    A podcast intro is a spoken sting, not a hook for a Reel and not a lesson. Open with the show name, name the listener, then one sentence of what this episode (or this show) is for. Spell numbers and brand names the way they should be heard. Time it out loud: most working intros sit between ten and twenty seconds. Write a matching outro at the same pass — thank you, where to subscribe, next week — so both files share one voice. If you wanted a 15–60s public hook, that is the TikTok / Reels voiceover how-to. If you wanted to radio-edit the interview, that is the Descript podcast page.

  2. 02

    Confirm a generated VO is the right format

    Use ElevenLabs when the missing piece is a short, realistic read you can regenerate when the season title changes. Skip it if the host already records a live open in the booth — film or record that, then cut it on the transcript. Skip it if the job is a long training narrator synced to slides. That is Murf Studio, and that how-to is an ordinary path on this site, not a hop here. Skip a library voice if the show is the host’s voice. That is a clone, and the full training-set path lives on the YouTube clone how-to.

  3. 03

    Open Text to Speech — the speech playground, not Agents

    ElevenLabs’ published product for this brief is Text to Speech: paste text, pick a voice, generate speech, download a file. That is the desk this page is written for. ElevenAgents is a different product (conversational agents). Studio is a longer project surface. We are not inventing a third workspace. This page sends you to ElevenLabs only.

  4. 04

    Sit on a plan that grants commercial rights before you treat the file as the show

    ElevenLabs’ own docs are explicit: you keep ownership of generated audio, but commercial usage rights come with paid plans. Free is an audition — published pricing lists 10,000 credits a month and no commercial license. Starter is the first paid tier that lists a commercial license and Instant Voice Cloning. Professional Voice Cloning is listed on Creator and above. Confirm the live grid on elevenlabs.io/pricing. We do not invent a checkout total or a commission rate. Do not publish a free-tier sting on a monetized feed.

  5. 05

    Pick a Voice Library voice — clone only if the show already is your voice

    Open Voices and browse Default Voices or the Voice Library. Preview before you apply. Cast one narrator and keep it for the intro and the outro. Filter toward narration / podcast-style reads, then listen on headphones. Clone only if listeners would notice a stranger. Instant Voice Cloning is the published self-serve path from short samples (ElevenLabs’ cloning help: about 1–2 minutes of clean audio) on paid plans that list it. Professional Voice Cloning trains a dedicated model on a longer set (published as 30–180 minutes) and requires Creator or above. Do not invent a clone wizard here — if you need the training-set workflow, use the clone how-to. Do not clone a guest from last week’s Riverside file.

  6. 06

    Paste the sting, pick a model, generate, then re-roll the take

    Type or paste the intro into the text box. Voice first, then model, then settings — that is the order ElevenLabs’ Text to Speech guide ranks. For a short English sting, Multilingual v2 is the published “most stable on long-form” model and still the safest default for a 15-second read you will reuse every week. Eleven v3 is the expressive model; it supports audio tags such as [sighs] or [clears throat], and it does not expose every older slider. Official starting settings for the sliders that exist: Stability around 50, Similarity around 75, Style exaggeration at 0. Speed defaults to 1.0 (range 0.7–1.2) where the model offers it. Spell out numbers. Then press Generate Speech. The model is nondeterministic — same text can yield a different take. Re-roll the open until the first three words punch; leave a clean body alone.

  7. 07

    Download the file, then drop it at the top of the episode in Descript or your editor

    After a generation, ElevenLabs’ help says you can download immediately from the control on the bottom right. Older takes live in History on the Text to Speech page — History lists MP3 (128 kbps) or WAV, with Advanced formats of MP3 192 / 256 kbps, M4A, and FLAC. Higher-quality options are listed on paid tiers; confirm the live download list. Name the file so the next recut is obvious: after-the-cut_intro_en_v1. Then put it on the episode: in Descript, import the sting into the same project as the interview and lay it before the conversation (Keep separate if the files are sequential — intro, then interview, then outro). In any other editor, drop the WAV or MP3 on a dedicated intro track, leave a breath before the host speaks, and duck any music bed under the VO. When the season name changes, edit the sentence, generate, bump to v2, replace the file. That recut is why you did not book a booth for a fifteen-second open.

Generated sting vs Murf narrator vs filming the host

“AI podcast intro” is a search, not a product. The decision is the format. The ElevenLabs column matches Text to Speech help and the live pricing page as of August 2026. Murf stays qualitative here — we already wrote that desk, and this page does not hop there.

Podcast open: ElevenLabs sting vs Murf narrator vs recording the host (August 2026)
Criterion ElevenLabs VOMurf studio narratorRecord the host
What the listener hears A short, realistic sting you can regenerate when the season title changes A directed studio narrator timed on a timeline — built for lessons, usable as a read The actual host, in the actual room, a file that ages when you rename the show
When it is the right buy You need a 10–20s intro/outro that sounds human and you will recut it The job is a longer directed read synced to slides or picture The open has to be that person, once, and will not be rewritten
What you re-do when the copy changes Edit the sentence, Generate Speech, replace the file on the timeline Edit the block, regenerate, replace the audio Re-book the host, the room, and an editor
Tool this desk writes for ElevenLabs Text to Speech — script, voice, Generate Speech, download Murf Studio — ordinary path: the training-VO how-to. Not a hop on this page A booth or a quiet room. Right when the voice has to be live
Best 2026 fit Weekly show open/close that should sound like a person, not a stock LMS read A module narrator you time to slides — see the Murf training page A one-time live open the host already records every week
The podcast-intro desk

ElevenLabs

A Text to Speech sting for a show that already has an episode. Free plan to audition a script; a paid plan is the commercial-rights download you can put on the feed.

The script is a sting, spoken

Do not write a cold open for a documentary. Do not write a first-second Reel hook. Time the copy out loud. Ten to twenty seconds is enough for the show name, who it is for, and one promise. Longer than that and you are writing the episode.

A working shape, spoken at a normal pace:

You’re listening to After the Cut — the weekly show for editors who still have to ship on Friday. I’m Maya Chen. Today: one messy timeline, one decision that stuck.

That is roughly a quarter-minute. Write the outro in the same sitting so both files share a voice:

That’s After the Cut. If this saved you an hour in the timeline, send it to the editor who is still in it. See you next week.

Spell the words the model should say. “Season three,” not “S3.” “Friday,” not “Fri.” ElevenLabs’ Text to Speech help is blunt about numbers and symbols: write them out, especially on multilingual models, because the same digit is pronounced differently across languages. A show called “Q3 ARR” needs those letters in the script the way you want them heard.

Keep one voice for the intro and the outro. A library narrator that survives your show name is worth more than a cinematic whisper that flubs the title. Emotional range is not why you are here. If realism-versus-timeline is the actual question, the comparison is ElevenLabs vs Murf vs Synthesys — ordinary path, no hop. Full product notes live in our ElevenLabs review.

Pick a voice. Generate the sting. Direct the take.

Open Text to Speech. Select a voice from the control ElevenLabs documents at the bottom left — Default Voices or the Voice Library. Preview. Apply one voice and keep it. A listener who hears v1 in March and v3 in October should still recognize the person who said the show name.

Clone only if that person has to be you. Instant Voice Cloning is the published fast path from short samples (about 1–2 minutes of clean audio in the cloning help). Professional Voice Cloning is the dedicated model: Creator plan or above, a longer training set (published as 30–180 minutes), and a wait while it fine-tunes. We are not going to invent a clone wizard on this page. If you do not already have the voice in My Voices, use How to clone your voice for YouTube and come back.

Paste the sting as one short block. Voice, then model, then settings — ElevenLabs ranks those in that order. Multilingual v2 is the published stable default and the one we would start on for a weekly English open. Eleven v3 is the expressive model: audio tags such as [sighs] or [clears throat], a 5,000-character cap, and fewer of the older sliders (Speed, Similarity, and Speaker Boost are documented as unavailable on v3). Flash models are the low-latency family; a fifteen-second preroll does not need 75ms.

Where the sliders exist, the official starting point is Stability around 50, Similarity around 75, Style exaggeration at 0. Speed defaults to 1.0; the documented range is 0.7 to 1.2. Lower Stability for a livelier take, then generate more than once — the model is nondeterministic. Higher Stability for a straighter announcer. Do not invent a slider we cannot see on the public page. If a control is missing on your model, change the sentence and re-roll.

Pauses: punctuation first. A dash or em-dash is the documented beat; ellipsis adds hesitation, which a show open usually does not want. On Multilingual v2, Flash v2, and Flash v2.5, ElevenLabs also documents an SSML break tag for a timed pause of up to three seconds. Confirm the live syntax in their Text to Speech help rather than pasting markup we cannot see unchanged. On Eleven v3, use audio tags and punctuation — that model’s help says it does not support SSML break tags.

Press Generate Speech. Listen on headphones. Re-roll the first three words until they punch. Leave a clean body alone. Two free regenerations of the exact same text and settings are the published allowance; any edit to the copy or the sliders is a new generation.

Download, then put the file where the episode already lives

The download is the gate. ElevenLabs’ Text to Speech help: after you generate, use the download control on the bottom right. Earlier takes sit in History on the same page — sidebar Text to Speech, then the history panel (or the history icon above Generate Speech on a narrow screen). History lists MP3 at 128 kbps or WAV; Advanced adds MP3 192 / 256 kbps, M4A, and FLAC. Higher-quality options are listed on paid tiers. For a podcast enclosure, WAV or a high-bitrate MP3 is the usual handoff. Confirm the live list on the plan you pay for.

Then hand the file to the episode you already have. In Descript, import the sting into the episode project and lay it before the conversation. Intro, interview, and outro are sequential files — keep them separate in the script rather than combining them as if they were host and guest ISOs of the same take. Studio Sound and filler removal are for the interview, not for a fifteen-second generated open. The working radio-edit path is How to edit a podcast in Descript. That page hops to Descript. This one does not.

In any other editor, put the sting on its own track. Leave a breath before the host speaks. If you use a music bed, duck it under the VO so the show name is the loudest thing in the first second. Name the export so the next person can find it: after-the-cut_intro_en_v1. When the season title changes, open the sentence, generate, bump to v2, replace the file. Leave the outro that is still true. That is the reason you did not film the host for the open.

Listeners still deserve a plain-language note when the open is synthetic. Say it in the show notes. If you also upload the episode as video, use the current platform disclosure for realistic synthetic or altered content. That is separate from the vendor license. You need both.

Commercial rights, the feed, and the next cut

Read the live page. ElevenLabs meters in character credits across the suite. Free is 10,000 credits a month and no commercial license. Starter is the first paid tier that lists a commercial license and Instant Voice Cloning — $6/month on the public grid we verified in August 2026 (elevenlabs.io/pricing). Creator is the tier that lists Professional Voice Cloning ($22/month on that same grid; annual billing is cheaper). Confirm checkout. We do not print a commission rate.

A weekly 15-second sting is a rounding error on credits. The plan question is the license, not the meter. Sit on a paid plan before the file is the one subscribers get. If you later want the host’s likeness on B-roll episodes, that is the clone how-to, not a reason to start this brief in Creator. If the remaining job is the interview, go back to the Descript radio edit. If the remaining job is a slide narrator, the Murf training page is that brief. The voice-tool comparison lives on ElevenLabs vs Murf vs Synthesys. This page does not hop there. The only Try button here is ElevenLabs.

Frequently Asked Questions

What is the best AI tool for a podcast intro in 2026? +
ElevenLabs Text to Speech, when the job is a 10–20 second intro or outro that has to sound human and you will regenerate it when the copy changes. Write the sting, pick a Voice Library voice (or an existing clone if the show is your voice), press Generate Speech, download MP3 or WAV, and drop the file at the top of the episode in Descript or your editor. Murf is the better pick when the missing piece is a directed studio narrator on a timeline — that is a training-VO job, and that how-to is an ordinary path on this site. Record the host if the open already happens live in the booth and will not be rewritten. Buying a slide-sync studio for a fifteen-second sting, or booking a half-day to rerecord “Season 3,” is the expensive mistake.
Can I make a podcast intro with ElevenLabs for free? +
You can audition. ElevenLabs’ published Free plan is 10,000 credits a month. That is enough to hear whether a library voice survives your show name. It is not a commercial license. ElevenLabs’ own docs: you retain ownership of generated audio, but commercial usage rights are only available with paid plans. Sit on a paid plan before the sting is the one subscribers hear. Confirm the live grid on elevenlabs.io/pricing. We do not invent a checkout total.
Should I clone my voice or pick a Voice Library voice for the intro? +
Pick a library voice unless listeners would notice a stranger. Most podcast intros are a show identity, not a host replica: one consistent narrator for the open and the close. Clone only if the show already is your voice. Instant Voice Cloning is the published self-serve path from about 1–2 minutes of clean audio on plans that list it. Professional Voice Cloning trains a dedicated model on a longer set (published as 30–180 minutes) and requires Creator or above. The full consent-and-training workflow lives at How to clone your voice for YouTube. Do not clone a guest from last week’s session.
When should I use ElevenLabs instead of Murf — or instead of filming the host? +
Use ElevenLabs when the deliverable is a short, realistic sting you will recut without a booth. Use Murf when the picture already exists — slides, a screen recording — and you need a directed studio narrator on a timeline. That is How to make training voiceovers with Murf, an ordinary site path, not a second money button. Film or record the host when the open is already a live performance and the words will not change. Same show calendar. Different file. Full voice-tool scorecard: ElevenLabs vs Murf vs Synthesys.
How do I add an ElevenLabs intro to Descript or another editor? +
Download the take from Text to Speech — immediately after Generate Speech, or later from History as MP3 or WAV. In Descript, start or open the episode project, import the sting, and lay it before the interview. If you drag intro and conversation in together, keep them sequential (Keep separate) rather than combining them into one sequence as if they were ISO mics of the same take. The radio-edit path for the interview itself is How to edit a podcast in Descript — ordinary page, no hop from here. In any other editor, put the file on its own track, leave a breath before the host speaks, and duck a music bed under the VO. Name the export so v2 is obvious.
Do I need to disclose that the podcast intro is AI-generated? +
Plan to, when a listener could take the voice as a real person speaking. That is separate from the vendor license. You need both: a paid ElevenLabs plan if the episode is commercial, and a plain-language note in the show notes or the current platform disclosure if you also publish the episode as video. YouTube Studio still asks about realistic altered or synthetic content on upload — use the live checkbox. An audio RSS host may not have a toggle; say it in the description anyway. Script-only AI that you then recorded yourself is a different case.
How is this different from the YouTube clone how-to and the Descript edit how-to? +
The clone page is a training-set and rights job: Instant vs Professional, consent, commercial-rights generation for channel narration. The Descript page is the radio edit of the episode: import, Studio Sound, filler review, export for YouTube or your RSS host. This page is the fifteen-second open (and the matching close): write the sting, pick or reuse a voice, generate, download, drop it on the timeline. Same company on the hop as other ElevenLabs pages. Different brief. If the file is the interview, start at the Descript how-to. If the file should sound like a named host and you do not have a clone yet, start at the clone how-to.

Continue the Pipeline

Sponsored

Try ElevenLabs