Skip to content
AI Video Tools Guide
Desk /
Menu
Guides
Consistent CharactersCinematic AI PromptsClone Your Voice for YouTubeClone Your Voice with ElevenLabsYouTube Voiceover with ElevenLabsYouTube Ad Voiceover with MurfAI Voiceovers for TikTok & ReelsTraining Voiceovers with MurfInstagram Ad Voiceover with MurfTikTok Voiceover with MurfPodcast Ad Voiceover with MurfPodcast Intro with ElevenLabsSales Voiceover with ElevenLabsLinkedIn Voiceover with ElevenLabsOnboarding Voiceover with ElevenLabsAdd AI Music to YouTube ShortsAdd AI Music to a YouTube ShortAdd AI Music to Instagram ReelsAdd AI Music to TikTokScore a YouTube Video with MubertAdd AI Music to a PodcastAdd AI Music to a Course TrailerAdd AI Music to a LinkedIn VideoTurn YouTube Videos into ShortsBatch-Clip a YouTube Channel with KlapRepurpose a Webinar into ShortsClip a Zoom Recording with VizardClip a Teams Meeting with VizardClip a Google Meet with VizardClip a Webinar with VizardClip a Podcast with VizardMake Podcast Clips with KlapLinkedIn Clips with KlapInstagram Clips with KlapTikTok Clips with KlapYouTube Shorts with KlapTranscribe a Podcast in DescriptEdit a Podcast in DescriptClean Up Podcast Audio in DescriptRemove Silence in DescriptAdd Captions in DescriptOverdub a Line in DescriptSplit Speakers in DescriptAdd AI Captions to YouTube ShortsMake an AI Avatar VideoAI Avatar Training VideoLocalize Training Videos with SynthesiaProduct Demo Videos with SynthesiaFaceless YouTube Channel with SynthesiaLinkedIn Videos with SynthesiaHR Onboarding Videos with SynthesiaSales Enablement Videos with SynthesiaCustomer Support Videos with SynthesiaCourse Trailer with SynthesiaExplainer Video with SynthesiaWebinar Recap with SynthesiaInternal Update with Synthesia
Guide · Editing & Post Verified August 2026

How to Split Speakers in Descript

A practical 2026 path for who said what: identify speakers after transcription, rename Speaker 1, fix a miss, and export a labeled transcript — or sequence the ISOs if you already have separate tracks. Built from Descript's published Detect, Speakers, sequence, and export help (August 2026).

By Scott /10 min read

Most “split speakers in Descript” posts promise two isolated tracks from a booth mix. Public help does not name that button. What they publish is a label pass on a mixed file — Detect, Identify speakers, rename — and a sequence pass when you already have separate mics. This guide is that working path, using the one editor we send for this brief: Descript.

When speaker labels are the job — and when you should transcribe, cut, or caption instead

Use this pass when the people are the deliverable. A weekly interview that still says Speaker 1, a guest who asked to review quotes under their name, a producer who will not open a file that attributes the wrong voice: you wait for detection, you Identify speakers, you rename, you export text. The same labels feed captions and a later radio edit. Those are later pages. This one stops at who said what.

Skip this pass if the script is empty. Getting the words is How to transcribe a podcast in Descript — an ordinary guide link, not a second money hop. Skip both if the remaining job is the ramble at 18:40: How to edit a podcast in Descript. Type on the picture is How to add captions in Descript. A flubbed line is How to overdub a line in Descript. Those pages hop to Descript too. This one already will.

The seven-step Descript speaker-label pass

  1. 01

    Confirm the interview is yours to label, then get a real file — mixed or ISOs

    This pass is for a podcast or interview you hosted, or a session you have written permission to transcribe and attribute. Guest likeness is not automatically yours to name in a public transcript. Put quote and speaker-label rights in the guest release. Then export a recording: a mixed WAV/MP3 from the booth, a video-podcast MP4, host and guest tracks from a recorder, or a Rooms session you captured in Descript. An RSS XML feed is not a speaker-label input. Prefer the file you recorded. If the master is not yours, stop.

  2. 02

    Start a project for this episode and put the file in the script

    Treat one episode as one project. Official getting-started copy gives you two starts: record in Descript, or import an existing file. Import paths they publish include upload from your computer, import from Zoom, and upload from a phone. Drag the file into the Script editor, or insert it from the Project panel. Official add-a-file help: only script media is transcribed. A file sitting as a layer is not a speaker map yet. You can also run Transcribe file from the options menu next to the file name. Every uploaded or recorded file counts against media minutes.

  3. 03

    Wait for the script — then Identify speakers when detection finishes

    Automatic speaker detection is on by default in App Settings. Official Detect and label speakers help: Descript detects voices, then you name them once. When detection finishes you get a notification; click Identify speakers. The Identify speakers modal opens. Add a name or pick an existing label. Listen to the samples. Close. Add the file to the script if it is not already there — the labels appear in the document. Pricing copy on descript.com/pricing names that clip-and-name assistant Speaker Detective. The help article names the follow-up Identify speakers. We write to the control you can click. If detection was off, open the file options menu in the Project panel and choose Detect speakers.

  4. 04

    Rename the labels so Speaker 1 is a person — @, Change speaker, or click any instance

    Official Speakers help: speakers are labels in the script. Type @, or select text and open the three-dot menu → Change speaker, then create a name or pick one you already used. Create speaker, use an existing label, or pick an AI Speaker if you are writing a new line — that last option is for generated speech, not for naming a recorded guest. Click any instance of a label to rename it everywhere. Official Replace in project with… swaps one label for another across the composition. Do not export a show-notes file that still says Speaker 2.

  5. 05

    Reposition a label that landed on the wrong voice — do not delete the paragraph

    Detection misses overlaps, similar voices, and a guest who laughs into the host mic. Official Speakers help: hover the speaker, click the reposition icon to the left, and drag the label to where that person actually starts. If a stretch is on the wrong name, highlight it and Change speaker / @. Removing a speaker label from the published options menu assigns those sections to the speaker above it; official copy says that does not delete the text or the audio. Fix the attribution. Leave the take.

  6. 06

    If you have separate tracks, Combine into sequence — a mixed file is not a one-click split

    Public Detect and Speakers help labels voices on a single audio or video file. They do not name a Split speakers or Split mixed track control that turns one booth mix into isolated host and guest WAVs. We describe the action they publish. If you already have ISOs — host mic, guest mic, Rooms tracks — drag those files into the script. Official Create a sequence help: assign speaker labels to each file when prompted, then Combine into sequence. Keep separate only if the files should sit end to end. Open the Sequence Editor (right-click the sequence, or Shift+Command+O / Shift+Ctrl+O) when one track needs mute, solo, or Remove from script / Include in script because of mic bleed. Official sequence cap is 14 tracks. Blade on the timeline splits a clip. It is not a speaker-split button.

  7. 07

    Export the labeled transcript — or solo a sequence track if you need an isolated file

    Official transcript export: click Export, open the Transcript tab. Published formats are .html, .md, .docx, .txt, and .rtf. Turn on speaker labels, or speakers in every paragraph, when the file is an interview. Timecode settings cover offset, interval, paragraph breaks, speaker labels, and markers. Wordless media and scene boundaries do not appear in the text file. If the remaining job is an isolated audio file from a sequence you already have, official Export individual sequence tracks as mono files: open the Sequence Editor, select the track, click S (solo) in the right sidebar, exit (Esc or Done), then File → Export. Descript exports the soloed track as a mono file. Repeat per track. That path needs a sequence. It does not invent isolated tracks from a mix. Then listen. If Speaker 2 is still in the sidecar, go back.

Mixed file vs sequence vs an isolated export

“Split speakers in Descript” is a search. The decision is what you already have. The columns match Detect / Identify speakers on one file, Combine into sequence on ISOs, and solo-then-export on a sequence, as of August 2026. We are not inventing a third product named Speaker Splitter — that is this same editor.

Descript speakers: mixed-file labels vs sequence ISOs vs soloed-track export (August 2026)
Criterion One mixed fileSeparate tracks (sequence)Need an isolated WAV
What you hand the tool One booth mix, a video-podcast MP4, or a single uploaded interview Host + guest ISOs, Rooms tracks, or recorder stems from the same conversation A sequence that already has one clip per speaker
Published control Detect speakers → Identify speakers. Then @ / Change speaker / rename any instance Create a sequence / Combine into sequence; assign a speaker label to each file when prompted Sequence Editor → S (solo) → File → Export. Official “export individual sequence tracks as mono files”
What “split speakers” actually means here Labels in the script — who said what on one file. Not isolated audio tracks One script driving aligned tracks. Include / exclude a track if bleed duplicates the guest A mono file of the soloed track. Repeat for the other speaker
When it is the right surface A weekly interview you recorded to one mix and need named in the transcript ISOs you still have and refuse to flatten before you label A mixer or NLE that wants host and guest as separate files
What public docs do not name A Split mixed track / Split speakers button that prints two WAVs from one mix A “always remember this guest” wizard beyond Detect / Identify / Create speaker A batch “export all speakers as stems” control we can confirm on a mixed file
Best 2026 fit Show notes, a guest review, or captions that should say Sarah not Speaker 2 Interview ISOs you will also radio-edit later A sequence you already built — not a mix you hoped would unbake
The speaker-label pass

Descript

Identify speakers, rename the labels, sequence ISOs if you have them, export a named transcript or a soloed track. Free is a short-episode demo; paid from $16/mo billed yearly.

Identify speakers — the published follow-up after transcription

Open a project for this episode. Official getting-started copy gives you two starts: record in Descript, or import an existing file. Import paths they publish include upload from your computer, import from Zoom, and upload from a phone. Rooms records guests remotely and keeps each participant on a separate track — useful when you already have ISOs. A mixed booth file is enough for labels. Separate stems are useful later if you will sequence. They are not required to name the voices.

Put the file in the script. Official add-a-file help: drag it from the Project panel into the Script editor, or select it and click Insert into script. Descript then transcribes the spoken audio and aligns each word. A file that is only a layer is not a speaker map. Official copy says so. If the script is empty after import, you added media to the canvas and not to the document. Insert it. Then wait.

Automatic speaker detection is on by default. Official Detect and label speakers help is written for a conversation in a single audio or video file. Detect: the AI finds voices. Identify: you assign a name. When detection finishes, click Identify speakers. Add a name or pick an existing one. Listen to the samples. Close. Labels follow the file into the script. Pricing copy on descript.com/pricing calls the clip-and-name assistant Speaker Detective and lists Detect 8+ speakers as a capability. The help article names the modal Identify speakers. We write to that click.

If detection was off, Detect speakers lives on the file options menu in the Project panel — the ellipsis next to the file. Official App Settings split automatic vs manual. Automatic (default) runs without asking how many speakers; published limit is up to 10 hours. Manual prompts for the count first; published limit is up to 3 hours. Always ask before detecting speakers is the toggle. If the file is longer than the published limit, Detect speakers is disabled. Official copy says you can check that by right-clicking the file. Split a multi-hour dump first. We are not inventing a “force detect on a 12-hour file” override.

Media minutes start when the file is in the project. Official review-era metering we already verified on the Descript review: every file you upload or transcribe counts. A one-hour interview plus a guest ISO is two files. Do not re-import an hour you already labeled “just to try detection again” unless you meant to spend the minutes.

Rename, reposition, then decide whether you even have tracks to isolate

After the labels exist, official Speakers help is the daily surface. Type @, or select text and Change speaker, then create a name or pick one you already used. Click any instance to rename the label everywhere. Replace in project with… is the published swap when you created “Host” and “Sarah” and later want them consistent. An AI Speaker on that menu is for generated speech — stock or a custom clone. It is not how you name a recorded guest.

Detection will miss. Two hosts who sound alike, a remote guest on a laptop speaker, a laugh that bleeds into the other mic: a paragraph lands on the wrong voice. Official fix: hover the speaker, click the reposition icon, drag the label to the first word that person actually said. Highlight a stretch and Change speaker if the miss is longer than a line. Removing a speaker from the published options menu assigns those sections to the speaker above it. Official copy: that does not delete the text or the audio. Use it to collapse a ghost label, not to cut the take.

A mixed file stops there. Public Detect and Speakers help do not name Split speakers, Split mixed track, or Separate stems. Blade on the timeline splits a clip so you can move a piece. It does not unbake a booth mix. If a YouTube tutorial shows duplicating the file and deleting the other person by hand, that is a workaround they invented. It is not a control we will print as official.

Sequence the ISOs if you have them — then export the file that left the app

If you recorded host and guest to separate tracks — Rooms, a recorder, a Zoom audio-record folder — drag those files into the script together. Official Create a sequence help: when prompted, assign speaker labels to each file, then Combine into sequence. Keep separate lays the files end to end. You can also select files in the Project panel, right-click, and choose Create sequence. The sequence shows up in the script and timeline as one item. Official help allows up to 14 tracks.

Open the Sequence Editor when one track is the problem. Official shortcuts: right-click the sequence, or Shift+Command+O (Mac) / Shift+Ctrl+O (Windows). Mute or solo from the right sidebar (speaker icon and S). Official Fix duplicate transcripts or mic bleed help: isolate the cleanest track per speaker. Select the track → ··· → Remove from script or Include in script so the guest is not printed twice. Align a late-start recorder before you trust the combined labels. We describe those actions. We are not inventing a “de-bleed this kitchen” button.

Then leave with the names, or with a file. Official transcript export: Export → Transcript tab. Formats are .html, .md, .docx, .txt, and .rtf. Turn on speaker labels when the file is an interview. Turn on speakers in every paragraph if your CMS wants a name on each block. Set timecode if a producer will jump to a quote. Official caveat: wordless media and scene boundaries do not appear in the text file. Read the export. If Speaker 2 is still in it, go back.

If the remaining job is an isolated audio file from a sequence you already built, official Export individual sequence tracks as mono files: Sequence Editor → select the track → S (solo) → Esc or Done → File → Export. Descript writes the soloed track as a mono file. Repeat for the other speaker. That path needs a sequence. It does not invent isolated tracks from a mix. Do not treat it as a caption sidecar — SRT and VTT live on Export → Subtitles, which is How to add captions in Descript.

Free-plan media minutes are 60 a month on the August 2026 grid. That is a short episode, not a season. Sit on a paid plan before the weekly hour is the file you label. Confirm checkout on descript.com/pricing. We do not print a commission rate.

When the next job is the transcript, the edit, captions, or a one-line replace

Stay here when you need the names. Hand the composition to the words pass when spellings are still wrong: How to transcribe a podcast in Descript. Hand it to the radio edit when a story still has to leave: How to edit a podcast in Descript. Hand it to captions when the picture needs type: How to add captions in Descript. Hand it to a one-line replace when the host misspoke: How to overdub a line in Descript. None of those are Try buttons on this page. The only hop here is Descript.

Finding 5–15 verticals in a finished episode is a clipper job — start at Descript vs Klap if you have not picked that stack yet. A licensed intro, outro, or bed is How to add AI music to a podcast. Klap and Mubert stay foils here, not hops.

We do not print an affiliate commission rate, and we do not invent a “this will earn you $X” line. The product question is whether the sidecar you export still says Speaker 2 — and whether you were honest that a mixed file does not unbake into stems.

Frequently Asked Questions

How do I split speakers in Descript in 2026? +
On a mixed podcast or interview, the published job is labels, not isolated stems. Import a real file into the script, wait for transcription, then Identify speakers when detection finishes (or Detect speakers from the file options menu if it was off). Rename each voice — @, Change speaker, or click any instance. Reposition a label that landed on the wrong paragraph. Export from Export → Transcript with speaker labels on. If you already have separate host and guest tracks, Combine into sequence and assign a speaker to each file. Public help does not name a Split speakers control that turns one mix into two WAVs. If you need an isolated file from a sequence, solo the track in the Sequence Editor and File → Export. If the remaining job is the words, the radio cut, type on the picture, or a one-line replace, those are the other Descript how-tos — ordinary site paths. This page hops to Descript.
How do I identify speakers after Descript transcribes an interview? +
Automatic speaker detection is on by default. Official Detect and label speakers help: when detection finishes, click Identify speakers. Name each voice or pick an existing label, listen to the samples, then Close. Add the file to the script if it is not already there. If detection was disabled, Project panel → file ellipsis → Detect speakers. App Settings has automatic vs manual detection: automatic runs without asking how many speakers (published limit up to 10 hours); manual prompts for the count first (published limit up to 3 hours). Always ask before detecting speakers is the toggle that switches those modes. If the file is longer than the published limit, Detect speakers is disabled — official copy says so. Split a multi-hour dump first. We are not inventing a “name my guest” wizard beyond those controls.
How do I rename Speaker 1 and Speaker 2 in Descript? +
Official Speakers help: click any instance of the label and rename it — the name follows the rest of the episode. Or type @ in the script, or select text → three-dot menu → Change speaker, then Create speaker or pick an existing label. Replace in project with… swaps one label for another across the composition. Hover a misplaced label, use the reposition icon, and drop it where that person actually starts. Do not export until Sarah is Sarah. An AI Speaker on that menu is for generated speech, not for renaming a recorded guest.
Can Descript split a mixed track into separate speaker audio files? +
Not as a named one-click control we can confirm in public help. Detect and label speakers is written for a conversation recorded in a single audio or video file — it labels the transcript. Create a sequence is written for files you already have as separate tracks. Export individual sequence tracks as mono files is written for a sequence: Sequence Editor → solo (S) → File → Export, then repeat. Blade splits a clip on the timeline. None of those articles name Split speakers or Split mixed track. If you only have a booth mix, label the script and export text. If you need isolated WAVs, record or export ISOs next time, or build a sequence from files you already have. We describe the action. We do not invent the button.
How do I export a Descript interview with speaker names — or isolate a sequence track? +
For the words: Export → Transcript tab. Official formats are .html, .md, .docx, .txt, and .rtf. Include speaker labels, or speakers in every paragraph, when the sidecar is show notes or a guest review. Timecode settings can stamp speaker labels too. Official caveat: wordless media and scene boundaries will not appear in the text file. For an isolated audio file from a sequence, official mono-export help: Sequence Editor → solo the track → exit → File → Export. That writes the soloed track as a mono file. It is not a caption sidecar — SRT/VTT is Export → Subtitles on the captions how-to. It is not the RSS enclosure unless you meant to publish one voice.
Can I label speakers in Descript for free? +
You can learn the workflow. The Free plan (descript.com/pricing, August 2026) is 60 media minutes a month, a one-time 100 AI credits, 720p with a watermark, and 5GB of storage. Media minutes count every file you upload or transcribe — a one-hour interview plus a guest ISO is two files. That is a short-episode demo, not a weekly hour. Hobbyist is $16/month billed yearly on the same verified grid ($24 month-to-month) for 10 media hours. Creator is $24/month yearly ($35 monthly). Confirm checkout. We do not print a commission rate.
How is this different from the Descript transcribe, podcast-edit, captions, and overdub how-tos? +
This page is who said what: Identify speakers, rename the labels, reposition a miss, sequence the ISOs if you have them, export a labeled transcript or a soloed sequence track. The transcribe page is the words — import, wait, Correct names, export text or stay in the script. The podcast-edit page is the radio cut — delete the ramble, review fillers, export for YouTube or RSS. The captions page is type on the picture — Captions layer, Wordbar timing, burned-in or SRT/VTT. The overdub page is one misspoken line — authorize a custom AI Speaker, Regenerate (formerly Overdub), preview the room. Same editor. Same /go/descript hop. Different brief.

Continue the Pipeline

Sponsored

Try Vizard