How to Clean Up Podcast Audio in Descript
A practical 2026 path for the audio pass: import the raw take, run Studio Sound, listen for artifacts, level the voices, and export a clean WAV or MP3. Built from Descript's published Studio Sound, auto-leveling, and audio-export help (August 2026) — plus when cleanup is enough versus a re-record or a separate voiceover tool.
Most “clean podcast audio with AI” posts skip the decision that actually saves the episode. The words are already right — or they are not. If they are, you need the fridge, the hallway slap, and the quiet guest to sit down, then a WAV or an MP3 you can hand to a host. If they are not, no enhancer will invent a take you did not record. This guide is the working path for the first case — import, Studio Sound, artifact check, level, export — using the one editor we send for this brief: Descript.
When Descript cleanup is enough — and when you should re-record
Use Descript when the take is the show. A weekly interview, a two-host table, a solo essay you already like: the voices should stay those people. Studio Sound is the published enhancer for spoken voice. Auto-leveling is the published loudness control. Together they are the cleanup pass, not the edit.
Skip this pass if the remaining job is the radio edit. Deleting the tangent at 18:40 is a transcript job. That working path is How to edit a podcast in Descript — an ordinary guide link, not a second money hop. Skip both if the file you actually need does not exist yet. A rewritten show open is the ElevenLabs intro how-to. A slide narrator is the Murf training page. ElevenLabs and Murf stay foils on this page, not Try buttons.
The seven-step podcast audio cleanup
- 01
Confirm the take is yours, then get the raw file — not a compressed feed enclosure
This pass is for a session you hosted or have written permission to process. Guest likeness and a licensed bed are not automatically yours to recut. Then get the actual recording: a booth WAV, host and guest ISOs from a recorder, or a Rooms capture. Do not start from an RSS enclosure you already published and lossily re-encoded. Cleanup cannot restore what a clipped analog stage or a 96 kbps re-download already threw away.
- 02
Start a project for this episode and import the raw take
Treat one episode as one project. Import from your computer, from Zoom if that is the source, or from a phone upload. You can also record straight into a composition, or remotely in Rooms, which keeps each participant on a separate track. Transcription starts after the file lands. Every uploaded or recorded file counts against media minutes — a one-hour interview plus a guest ISO is two files against the same pool. Wait for the script before you start toggling effects; you need to hear the same words the listener will hear.
- 03
Name the problem before you enhance anything
Listen once on headphones, untreated. Steady HVAC or fridge hum, room slap, keyboard clicks, and two voices at different distances are cleanup jobs. A take that clipped in the recorder, a guest who is unintelligible, or a line you wish you had said differently is not. Studio Sound is built for spoken voice — official help describes it as reducing background noise, echo, and other distractions. It is not a mastering suite and it is not a rewrite. If the words are wrong, stop here and decide whether to re-record or replace that line.
- 04
Apply Studio Sound, then back the Intensity off until the voice still sounds like a person
Select the script layer. Open the Properties panel in the right-hand sidebar. In Audio Effects, toggle Studio Sound on. Official help also lists it under Sound Good in the AI Tools panel, and you can apply it on a selected track in the Sequence Editor if only one mic is the noisy one. It is a file-level effect: once on, it follows that file everywhere it appears. It needs a network connection and it burns AI credits. Use the Intensity slider to control how much enhancement is applied. Start lighter than the default if the room was already treated. Do not stack it on a file you already ran through a different enhancer without checking for metallic artifacts.
- 05
Level the voices, then listen specifically for artifacts
Studio Sound is the cleanup. Leveling is a separate published control. Auto-leveling adjusts clip loudness toward a standard level (around −16 dB). Toggle Auto-level on a clip from the Audio section of the Properties panel, or turn Perform automatic leveling on clips on for the project. Then listen on headphones with Studio Sound on and off. Official troubleshooting is blunt: if the waveform goes flat or silent, the noise was loud enough that the effect ate the speech — reduce Intensity or disable it for that file. Swallowed consonants, a watery S, a gated breath, or a voice that suddenly sounds like a phone codec are the stop signs. A known Chrome export bug can also add echo or drift after you leave the editor; Descript publishes Repair audio drift for that encode error. Do not export yet if the treated take sounds worse than the raw one.
- 06
Decide: keep the cleanup, re-record the take, or replace a line
Keep the Descript pass when the words are right and the remaining problem is steady noise, room, or uneven loudness. Re-record when the take clipped, the guest is unintelligible, or Intensity has to sit so high that the voice no longer sounds human. Replace a single line only when the host would rather generate that sentence than book the room again — that is a voiceover job, not a cleanup job. The working paths on this site are the ElevenLabs podcast-intro how-to for a short sting and the Murf training how-to for a directed narrator. Those are ordinary pages, not hops from here. Do not run Studio Sound on an AI voice clip until you convert it to an audio layer; official help says Studio Sound will not apply to generated speech until you do.
- 07
Export a clean WAV or MP3 you can hand to a host or an editor
Open Export in the top right. Choose the Audio tab. Set the destination to Local export. Published formats are .m4a, .wav, and .mp3. For a cleanup master you may still recut, WAV is the honest handoff. For an enclosure your RSS host will take, MP3 or M4A is the usual file — pick what the host accepts. Official export settings include channels, sample rate, bitrate (M4A/MP3), and Normalize volume with LUFS targets including −16 and −18. Fill metadata when the format supports it; official help notes that .wav does not carry artwork or chapter markers. Listen to the exported file, not only the editor preview. If this episode also still needs the radio edit — delete the ramble, review fillers — that is the Descript podcast-edit how-to, an ordinary site path.
Cleanup vs re-record vs replacing the line
“AI podcast cleanup” is a search. The decision is whether the take survives. The Descript column matches Studio Sound, Auto-leveling, and the live audio-export page as of August 2026. Voiceover tools stay qualitative here — we already wrote those desks, and this page does not hop there.
| Criterion | Descript cleanup | Re-record the take | Replace the line (VO tool) |
|---|---|---|---|
| What you hand the tool | A raw spoken take you still want to keep | A room, a mic, and the same script again | A sentence that should not stay as the recorded voice |
| What actually gets fixed | Steady noise, room, echo, and uneven clip loudness on spoken voice | Clipping, a ruined take, or a line you never said well | A missing or rewritten line — not HVAC, not slap |
| When it is the right move | The words are right. The room is the problem. | The words are wrong, clipped, or unintelligible at any Intensity | You would rather generate one sentence than book the booth |
| Published controls this desk writes for | Studio Sound + Intensity; Auto-level; Export → Local export as WAV/MP3/M4A | A quiet room. Then import the new file and treat it once. | A separate voice desk — ElevenLabs intro or Murf training. Ordinary paths, no hop here |
| Best 2026 fit | Weekly spoken episode that should stay the real hosts | A take you cannot ethically or sonically publish | A sting, a retake you will not schedule, or a narrator that is not this show |
Descript
Import the raw take, apply Studio Sound, check artifacts, level the clips, export WAV or MP3. Free is a watermarked demo; paid from $16/mo billed yearly.
Import the raw take — then name the noise
Open a project for this episode. Official getting-started copy gives you two starts: record in Descript, or import an existing file. Import paths they publish include upload from your computer, import from Zoom, and upload from a phone. Rooms records guests remotely and keeps each participant on a separate track. That matters later: Studio Sound can sit on one track in the Sequence Editor when only the kitchen-mic guest is the problem.
Bring the highest-quality file you still have. A session WAV from the booth or the recorder is the right input. A mixed program file is fine if that is all you were handed. An RSS enclosure you already posted is a last resort — you are cleaning a compressed copy of a compressed copy. Media minutes start when the file is in the project. Plan the import before you drag a camera card “just in case.”
Then listen once with nothing on. Write down what you actually hear. A steady air-conditioner or fridge is the job Studio Sound is built for. Room slap on a hard kitchen is the same family — official copy names background noise, echo, and reverb on spoken voice. Two hosts at different distances is a leveling job. A take that slammed the converter, a guest you cannot understand, or a sentence you wish you had not said is a different brief. Write that down before you spend credits.
Studio Sound for hum, room, and the rest of the noise floor
Studio Sound is the cleanup name on the marketing site and in help. Official language: an AI-powered audio effect that enhances spoken voice by reducing background noise, echo, and other distractions. Descript’s product pages also describe common spoken-voice distractions — AC hum, room echo, hiss, keyboard clicks — as the kind of problem the effect is aimed at. We are not going to invent a toolbar button called Hum Removal or Room Noise. The published action is: toggle Studio Sound, then move Intensity.
Apply it from the Properties panel in a composition, or on a selected track in the Sequence Editor, or from Sound Good in the AI Tools panel. Official help is explicit that it is file-level: one toggle follows every instance of that file. It wants a network connection. It burns AI credits. A treated booth often wants less than a default pass. A laptop in a kitchen is why the feature exists. Do not assume it will save a file that clipped in the recorder.
Intensity is the published amount control. Official troubleshooting: if the waveform goes flat or silent, the recording likely has very loud background noise and Studio Sound may be suppressing the speech. Reduce Intensity or disable the effect for that file. That is the first artifact check, not a later one. Listen on headphones. Toggle the effect off and on over the same sentence. If the consonants get duller before the hum is gone, you are past the useful range. Leave it there, or stop.
Underlord can be asked to enhance the audio. Treat that like the same file-level pass, then watch it. Full Underlord access is a plan question on the August 2026 grid; we do not invent a checkout total. The Free plan’s AI credits are one-time. Studio Sound spends them. Confirm the live page on descript.com/pricing before you treat Hobbyist as a weekly cleanup plan.
Level the take, then hunt artifacts on purpose
Leveling is not a second Studio Sound slider. Auto-leveling is the published control that keeps clip loudness consistent — official help says it targets around −16 dB. Toggle Auto-level on a clip from the Audio section of the Properties panel. For a whole project, File → Project settings → Perform automatic leveling on clips. For a default across projects, App settings → Automatic volume levels. You can also apply Auto-level to a clip inside the Sequence Editor when one ISO is the quiet one.
Official troubleshooting for a clip that will not come up: a brief peak — a sneeze, a bang — can throw the leveler off. Split around the spike, then re-level the quieter portion. We are describing that action. We are not inventing a “de-ess” or “gate” knob that public Studio Sound help does not name. If you later need a mastered mix with adaptive speaker leveling, Descript publishes Mix Audio / advanced mixing settings as a different surface. This page is the spoken-voice cleanup pass, not that mastering desk.
Now listen for what the enhancer traded away. Swallowed plosives. An S that turns watery. A breath that gates. A voice that suddenly sounds like a narrow phone codec. Those are stop signs. Drop Intensity. If the treated file is worse than the raw one, turn Studio Sound off. A known Chrome bug (macOS and Windows) can add echo or delay on local exports that used Studio Sound; Descript publishes Repair audio drift for that encode error. Check the file that left the app, not only the editor preview.
Music beds and sound effects are not what Studio Sound is tuned for. Official product copy is explicit that it is a spoken-voice enhancer. Do not run it on the sting you already generated. Official help: Studio Sound will not apply to an AI voice clip until you convert that clip to an audio layer. Even then, think twice — you are re-processing a voice that was never in a room.
Export a clean WAV or MP3 — then stop, or keep cutting
The cleanup is not done until a file leaves Descript. Official export overview: click Export, pick a destination. For a file you control, set Destination to Local export and open the Audio tab. Published formats are .m4a, .wav, and .mp3.
If another editor still has to touch this episode, export WAV. If the next hop is an RSS host, export MP3 or M4A — Apple’s public guidance still names MP3 or AAC; pick what your host lists. Official advanced settings include mono or stereo, 44.1 or 48 kHz, bitrate on M4A and MP3, and Normalize volume (Off, Peak, and LUFS targets including −14, −16, −18, −23, and −24). Use −16 LUFS when you want the same target Auto-leveling already aims at. Fill metadata — show title, episode title, description, artwork, chapters — when the format supports it. Official help says .wav does not carry artwork or chapter markers.
Descript does not mint your show’s RSS URL. Upload the enclosure to the host that already serves the feed. We do not invent a host hop. If this episode still needs the radio edit, do that before you treat the enclosure as finished: How to edit a podcast in Descript. That page hops to Descript too. This one already did. There is no second money button here.
Free-plan exports are 720p with a watermark on video. Do not publish a watermarked demo to a feed you already monetize. Sit on a paid plan before the cleaned file is the one subscribers get. Confirm checkout on descript.com/pricing. We do not print a commission rate.
When to re-record, and when a VO tool is the honest next hop
Re-record when cleanup is arguing with physics. A clipped converter. A guest under a leaf blower. Intensity so high the host no longer sounds like a person. Official help already tells you the silent-waveform case is “too loud to separate.” Believe it. Book the room, or the next quiet hour, and import the new file. Treat that file once.
Replace a line when the words are the problem and you will not get the host back. That is not Studio Sound. A short show identity sting lives on How to make a podcast intro with ElevenLabs. A directed narrator on picture lives on How to make training voiceovers with Murf. The voice-tool scorecard is ElevenLabs vs Murf vs Synthesys. None of those names are hops on this page. The only Try button here is Descript.
After the audio sits, the remaining weekly jobs are usually the radio edit and, later, clips. The edit path is the Descript how-to above. Finding 5–15 verticals in a finished episode is a clipper job — start at Descript vs Klap if you have not picked that stack yet. Klap stays a foil here, not a Try button.
We do not print an affiliate commission rate, and we do not invent a “this will earn you $X” line. The product question is whether the treated take survives a headphone pass after Studio Sound and Auto-level — and whether you were honest about the takes that will not.
Frequently Asked Questions
How do I clean up podcast audio in Descript in 2026? +
Does Descript Studio Sound remove hum and room noise? +
How do I level podcast audio in Descript? +
When is Studio Sound not enough — should I re-record? +
When should I use a voiceover tool instead of cleaning the take in Descript? +
How do I export a clean WAV or MP3 from Descript? +
How is this different from the Descript podcast-edit how-to? +
Continue the Pipeline
- Guide How to edit a podcast in Descript →
- Comparison Descript vs Klap: transcript editor or clipper? →
- Guide How to make a podcast intro with ElevenLabs →
- Guide Add AI music to a podcast →
- Guide How to make training voiceovers with Murf →
- Review Descript review: transcript editor, meters & Underlord →
- Comparison ElevenLabs vs Murf vs Synthesys →