How to Clone Your Voice with ElevenLabs
A practical 2026 path: record or upload your own samples, start Instant Voice Cloning or Professional Voice Cloning, generate a test line in Text to Speech, and drop the take on a short voiceover. Built from ElevenLabs' live voice-cloning help and pricing (August 2026).
Most “clone your voice with ElevenLabs” posts skip the two facts that actually decide the replica. First: Instant Voice Cloning and Professional Voice Cloning are different products — different sample lengths, different waits, different plan gates. Second: a free-tier preview is not a voice you put under ads. This guide is the working path for the training set — consent, samples, Instant or Professional, a test line, then a short VO — using the one voice desk we send for this brief: ElevenLabs.
When you need a clone — and when a library voice is enough
Clone when the file has to be you. A host who already talks on camera and wants the same timbre on B-roll weeks. A founder whose buyers would notice a stranger on the pricing-page loop. A podcast that is one named voice. The economic case is the recut: edit the sentence, generate again, replace the clip. You do not re-book a booth every time a number changes.
Skip the clone if nobody would notice a stranger. A faceless explainer, a listicle, a demo bed with no existing vocal identity — pick a Voice Library narrator on the matching how-to and generate the read. That is faster than a training set. Skip both if you already record the line yourself and the words will not change. That file is a recording, then a cut.
The seven-step ElevenLabs voice clone
- 01
Confirm you own the voice — consent is the start, not a footer
Clone a voice you own, or stop. Instant Voice Cloning asks you to confirm that you have the right and consent before you save the voice. Professional Voice Cloning is stricter: ElevenLabs’ published PVC help says you can only create a Professional Voice Clone of your own voice, and a verification step is required. Even with written consent, you cannot submit someone else’s voice as a PVC on your account — they can train and verify it on theirs, then share it. Do not upload a guest, a commenter, a celebrity, or last week’s Zoom. Platform terms and replica laws are not theoretical.
- 02
Decide Instant Voice Cloning or Professional Voice Cloning
ElevenLabs names two products, not one slider. Instant Voice Cloning (IVC) is the fast path: about 1–2 minutes of clean audio, ready as soon as you save it. It does not train a dedicated model. Professional Voice Cloning (PVC) fine-tunes a dedicated model on a longer set — published as 30–180 minutes, with “at least an hour, closer to two or three” as the quality target — and usually takes 3–6 hours (sometimes longer in queue). Use Instant to hear whether the likeness is even in the neighborhood. Use Professional when the file will be a named person’s voice on a published VO. A Voice Library narrator is a different job — that is a stock read, not a clone.
- 03
Sit on a plan that lists the clone type and commercial rights
Instant Voice Cloning and a commercial license are listed on Starter and above. Professional Voice Cloning is listed on Creator and above; Free and Starter publish no PVC slots. ElevenLabs’ own billing language: you keep ownership of generated audio, but commercial usage rights come with paid plans. Free is 10,000 credits a month and an audition — not a clone you put on a monetized video or a landing page. Confirm the live grid on elevenlabs.io/pricing. We do not invent a checkout total or a commission rate.
- 04
Record or upload clean samples — capture quality beats file type
Record in a quiet, deadened room: one speaker, no music bed, no reverb, no crosstalk. Keep tone, accent, and pacing consistent — the model copies breathing, “um”s, and room noise as readily as timbre. ElevenLabs recommends MP3 at 192 kbps or higher; uncompressed WAV usually does not improve the clone and can stall the upload. Aim for a balanced level (published target: about −23 to −18 dB RMS, true peak around −3 dB). For Instant, about one to two minutes is the published sweet spot — more than about three minutes is documented as little help and sometimes worse. For Professional, the published floor is thirty minutes; more clean runtime is better. Include the words you actually say: product names, numbers, abbreviations. If you do not already have files, both clone flows let you record in the interface.
- 05
Start the clone from Voices — Instant or Professional, then wait if you must
Open Voices in the ElevenLabs dashboard. For Instant: use the plus control, choose Instant Voice Clone, upload or record, name the voice, confirm consent, then save. The clone is usable as soon as it appears under My Voices. For Professional: choose Create Voice, then Professional Voice Clone. Upload samples or record yourself. Use the published Audio settings control when a clip needs noise reduced or a second speaker pulled out. Complete the voice-verification prompt with similar gear and delivery to the samples — if it fails, the help says you can wait 24 hours or contact support. Then wait for fine-tuning. Status lives on the voice in My Voices / Personal; you also get an email when it is ready. Do not invent a third wizard. If a label on your account differs from these published names, follow the on-screen Instant Voice Clone or Professional Voice Clone action.
- 06
Generate a test line in Text to Speech before you trust the clone
Open Text to Speech — the speech playground, not Agents. Apply the new voice from My Voices (Instant) or the Personal tab (Professional). Paste one real sentence you would actually publish: a name, a number, a product. Voice first, then model, then settings — that is the order ElevenLabs’ Text to Speech guide ranks. Multilingual v2 is the published stable default for a short English line. Official starting points where the sliders exist: Stability around 50, Similarity around 75, Style exaggeration at 0, Speed at 1.0 (documented range 0.7–1.2). Press Generate Speech. Listen on headphones. Re-roll the line until the first three words sound like you. A demo sentence that never appears on the channel is a poor test. Two free regenerations of the exact same text and settings are the published allowance; any edit is a new generation.
- 07
Use the clone on a short VO, then disclose the replica
When the test line holds up, paste the short voiceover — a 10–20 second sting, a 15–45 second pitch, or the open of a longer script — and generate again. Download from the control ElevenLabs documents on the bottom right, or pull the take from History as MP3 or WAV. Name the file so the next recut is obvious:
scott_clone_test_en_v1. Drop it on a dedicated VO track, leave a breath before picture or a CTA card, and duck any music bed under the speech. When a fact changes, edit the sentence, generate, bump to v2, replace the clip. That recut is why you trained the clone. Disclose a realistic replica in the description or show notes, and use the current platform checkbox when you upload video. That note is separate from the vendor license. You need both.
Instant vs Professional vs a stock library voice
“Voice cloning” is a search, not a product. The decision is Instant, Professional, or no clone at all. The columns below match ElevenLabs’ published cloning help and the live pricing page as of August 2026.
| Criterion | Instant Voice Cloning | Professional Voice Cloning | Voice Library (no clone) |
|---|---|---|---|
| What it is | A fast likeness from a short sample — no dedicated model is trained | A fine-tuned model of your voice on a longer training set | A stock narrator licensed on the plan you generate from |
| Published sample target | About 1–2 minutes of clean audio; more than ~3 minutes is documented as little help | 30–180 minutes; docs push an hour, closer to two or three, for the best result | None — you pick a voice, you do not upload yourself |
| When it is ready | As soon as you save it under My Voices | After verification and fine-tuning — usually 3–6 hours, email when ready | Immediately after you apply the voice |
| Who you may clone | A voice you own or have consent to clone — you confirm that on save | Your own voice only — verification is required; you cannot PVC someone else | Not a clone. Do not treat a library voice as a named person |
| Plan the public grid lists | Starter and above (with a commercial license on paid plans) | Creator and above — Free and Starter list no PVC slots | Text to Speech on Free to audition; commercial rights on a paid plan |
| Best 2026 fit | A same-day audition of whether the likeness is even close | A published VO that has to be you — channel, podcast, or sales close | A faceless read when nobody would notice a stranger |
ElevenLabs
Instant Voice Cloning on a paid plan that lists it; Professional Voice Cloning on Creator and above. Free plan to audition a line; a paid plan is the commercial-rights download.
Record the samples like they are the product
The clone cannot outrun the source. Instant Voice Cloning’s help is blunt: about one to two minutes of clear audio, no reverb, no artifacts, no background noise. More than about three minutes is documented as little improvement and, in some cases, worse. How the room was captured matters more than how many files you attach. Professional Voice Cloning’s help is equally blunt in the other direction: thirty minutes is the floor; an hour, closer to two or three, is the quality target. Split long PVC sessions into roughly half-hour files if you are uploading hours.
One quiet room. One mic position. One speaker. No music bed. Keep the performance consistent — animated throughout or subdued throughout, not a mix. If the published VO is a calm narrator, do not train on a shouted hook. If the channel says “Q3 ARR” and “ChatGPT,” put those phrases in the set. ElevenLabs recommends MP3 at 192 kbps or higher. WAV usually does not make a better clone and can stall the upload. Published level target: about −23 to −18 dB RMS, true peak around −3 dB.
If you do not already have files, record in the interface. Instant Voice Cloning’s guide: upload or record, then follow the on-screen prompts. Professional Voice Cloning’s guide: Upload samples or Record yourself, with sample scripts for narrative, conversational, and advertising reads if you need copy. Use Audio settings on a PVC clip when the published controls for noise or a second speaker would clean the take. Do not invent a button that is not on your screen — describe the action (upload, record, reduce noise, verify) and use the live label.
Start Instant or Professional. Then generate a test line.
Instant: Voices, plus control, Instant Voice Clone. Upload or record. Name the voice. Confirm you have the right and consent. Save. Open My Voices and use it.
Professional: Voices, Create Voice, then Professional Voice Clone. Upload or record. Check the sample-length feedback. Process a dirty clip if the Audio settings control is there. Verify that the voice is yours, with similar gear and delivery. Then wait. Fine-tuning is usually 3–6 hours; the help says it can run longer in queue, and you get an email when the model is ready. Using a PVC before that finish is the documented “No model found for this voice” error. When it is ready, the Personal tab is the published place to use it.
A clone you have not heard on a real sentence is not a clone you should publish. Open Text to Speech. Apply the new voice. Paste one line you would actually ship — including a name and a number:
I’m Scott. This is the August recut of the retention graph — eight seconds, not a thumbnail pack.
Voice, then model, then settings. Multilingual v2 is the published
stable default for a short English test. Eleven v3 is the expressive
model (audio tags such as [sighs]; fewer of the older
sliders). Where the sliders exist, start at Stability around 50,
Similarity around 75, Style exaggeration at 0, Speed at 1.0. Press
Generate Speech. Listen on headphones. Re-roll the
first three words. Leave a clean take alone. If Instant misses the
accent, change the samples or step up to Professional — you cannot
retune accent after the fact.
Put the clone on a short voiceover
The test line is the gate. If it holds up, paste the short VO you actually need. A podcast sting. A sales close. The open of a YouTube explainer. Generate that block as its own take so a flubbed sentence does not burn a clean close. Download from the control on the bottom right, or from History — MP3 at 128 kbps or WAV, with higher-quality options listed on paid tiers. Confirm the live download list.
Import the file onto a dedicated VO track. Line the first word to
picture. Leave a breath before a CTA card. Duck a music bed under
the speech. Name the export so v2 is obvious:
scott_clone_open_en_v1. When a fact changes, edit that
sentence, generate, replace the clip. Leave the lines that are still
true. That is the reason you trained a clone instead of booking the
room again.
The finished brief still lives on the script pages. A 10–20 second show open is the podcast-intro how-to. A 2–5 minute narration is the YouTube-voiceover how-to. A 15–45 second pitch is the sales-voiceover how-to. Those pages do not hop from here. They assume the voice already exists. This page is how the voice gets into My Voices.
Listeners still deserve a plain-language note when the voice is a realistic replica. Say it in the description or the show notes. If you upload video, use the current platform disclosure for realistic synthetic or altered content. That is separate from the vendor license. You need both.
Commercial rights, the replica, and the next cut
Read the live page. ElevenLabs meters in character credits across the suite. Free is 10,000 credits a month and no commercial license. Starter is the first paid tier that lists a commercial license and Instant Voice Cloning — $6/month on the public grid we verified in August 2026 (elevenlabs.io/pricing). Creator is the tier that lists Professional Voice Cloning ($22/month on that same grid; annual billing is cheaper). Confirm checkout. We do not print a commission rate.
A thirty-second test line is a rounding error on credits. The plan question is the license and the clone type, not the meter. Sit on a paid plan before the file is the one subscribers or paid traffic hear. If Instant is close enough for drafts and Professional is the publish voice, that is a Creator question, not a reason to start in a higher tier. If the remaining job is the script, go back to the podcast-intro, YouTube-voiceover, or sales-voiceover how-to. The voice-tool comparison lives on ElevenLabs vs Murf vs Synthesys. This page does not hop there. The only Try button here is ElevenLabs.