How to Make an Audiobook Sample with ElevenLabs
A practical 2026 path: write a 60–90 second first-page hook, pick a Voice Library voice, generate the read in Text to Speech, download WAV, and drop it in the retailer or Kickstarter sample slot. Built from ElevenLabs' live Text to Speech help and pricing (August 2026).
Most “AI audiobook” posts skip the two facts that actually decide the file. First: a seventy-second first page is not a sales close, and it is not the whole book. Second: a free-tier preview is not a track you put on a retailer listing. This guide is the working path for the spoken sample — one opening scene, one voice, generate, export, drop it in the sample slot — using the one voice desk we send for this brief: ElevenLabs.
When a stock ElevenLabs voice is enough — and when you want a booth or a lesson desk
Use ElevenLabs when the open is a short, literary read you will regenerate. The first page of a novel on a store player. A Kickstarter audio block under the title. An author-site embed that still sounds like the person who said the place name last month. The economic case is the recut: edit the sentence, generate again, replace the file. You do not re-book a booth for seventy seconds.
Skip ElevenLabs if the author already records a live sample and the words will not change. That file is a recording, then a cut. Skip both if the remaining job is a scored lesson on slides. That is How to make a course lesson voiceover with Murf — an ordinary guide link, not a second money hop. Murf is the foil, not a Try button. Skip a stock library voice if the book is a named author’s voice and you do not have a clone yet. That is How to clone your voice with ElevenLabs — also an ordinary path from this page. Skip this brief if you are producing the whole title in chapters. Studio is the published longer surface; this page stays on the speech playground.
The seven-step ElevenLabs audiobook sample
- 01
Write a 60–90 second first-page hook — scene, one name, one unfinished question
An audiobook sample is a spoken first page, not a landing-page close and not a fifteen-second show sting. Time it out loud: sixty seconds is the opening scene plus one name; ninety is that plus the line that makes a listener stay for chapter one. Open in the room, not with a product. Name the person the listener will follow. Stop before you explain the plot. Spell place names, years, and titles the way they should be heard on a retailer sample player. If you wanted a 15–45s sales pitch, that is the sales-voiceover how-to. If you wanted a 10–20s show open, that is the podcast-intro how-to.
- 02
Confirm a generated sample is the right spoken-book file
Use ElevenLabs when the missing piece is a short, literary read you can regenerate when the first paragraph or the title page changes, and the sample slot already exists — a retailer player, a Kickstarter audio block, a BookFunnel preview, an author-site embed. Skip it if the remaining job is the whole book. That is chaptered work, and ElevenLabs Studio is a different surface — we are not inventing a full-title pipeline on this page. Skip it if the file is a scored lesson timed to slides. That is Murf Studio, and that how-to is an ordinary path on this site. Skip a library voice if the book is a named author and buyers would notice a stranger. That is a clone, not this brief.
- 03
Open Text to Speech — the speech playground, not Agents or Studio
ElevenLabs’ published product for this brief is Text to Speech: paste the first page, pick a voice, generate speech, download a file. That is the desk this page is written for. ElevenAgents is a different product (conversational agents). Studio is a longer project surface for chaptered work — right for a title you will produce in parts, wrong for a ninety-second retailer hook you will recut twice this month. Image & Video on the pricing page is a different surface — we are not inventing an MP4 download from the speech playground. This page sends you to ElevenLabs only.
- 04
Sit on a plan that grants commercial rights before the sample is the store listing
ElevenLabs’ own docs are explicit: you keep ownership of generated audio, but commercial usage rights come with paid plans. Free is an audition — published pricing lists 10,000 credits a month and no commercial license. Starter is the first paid tier that lists a commercial license and Instant Voice Cloning. Professional Voice Cloning is listed on Creator and above. Confirm the live grid on elevenlabs.io/pricing. We do not invent a checkout total or a commission rate. Do not publish a free-tier read on a retailer sample, a Kickstarter that sells the book, or a paid BookFunnel drop.
- 05
Pick a Voice Library voice — one narrator for the sample and the later recut
Open Voices and browse Default Voices or the Voice Library. Preview before you apply. Cast one narrator and keep it for this sample and for the recut when the first paragraph or the subtitle changes. Filter toward narration-style reads, then listen on headphones and on a phone speaker — that is how most sample players sound. A spoken-book hook should sound like a reader finishing a thought, not a cinematic whisper and not a first-second Reel claim. A stock voice is enough when the title has no existing vocal identity. Clone only if listeners would notice a stranger. Instant Voice Cloning is the published self-serve path from short samples (ElevenLabs’ cloning help: about 1–2 minutes of clean audio) on paid plans that list it. Professional Voice Cloning trains a dedicated model on a longer set (published as 30–180 minutes) and requires Creator or above. Do not invent a clone wizard here. Do not clone a narrator from someone else’s Audible page.
- 06
Paste the 60–90s page, pick a model, generate, then leave room after the last line
Type or paste the first page into the text box as one short block. Voice first, then model, then settings — that is the order ElevenLabs’ Text to Speech guide ranks. For a short English literary read, Multilingual v2 is the published “most stable on long-form” model and still the safest default for a line you will reuse on a store listing. Eleven v3 is the expressive model; it supports audio tags such as [sighs] or [clears throat], and it does not expose every older slider. Official starting settings for the sliders that exist: Stability around 50, Similarity around 75, Style exaggeration at 0. Speed defaults to 1.0 (range 0.7–1.2) where the model offers it — nudge Speed down, or add a dash / em-dash, when the last sentence runs into the silence. Spell out numbers and years. Then press Generate Speech. The model is nondeterministic — same text can yield a different take. Re-roll the first sentence until the room is clear; leave a clean last line alone.
- 07
Download WAV (or high-bitrate MP3), then drop it in the sample slot you already have
After a generation, ElevenLabs’ help says you can download immediately from the control on the bottom right. Older takes live in History on the Text to Speech page — History lists MP3 (128 kbps) or WAV, with Advanced formats of MP3 192 / 256 kbps, M4A, and FLAC. Higher-quality options are listed on paid tiers; confirm the live download list. Text to Speech is an audio export. We do not invent an MP4 button on that playground. Name the file so the next first-page recut is obvious:
salt-line_sample_en_v1. Then put the audio where a listener already expects a preview: the retailer sample player, the Kickstarter audio block, the BookFunnel or author-site embed. Confirm the live upload rules on the store you use — we do not invent an ACX or Findaway button inside ElevenLabs. When the opening paragraph or the title page changes, edit the sentence, generate, bump to v2, replace the file. That recut is why you did not book a booth for seventy seconds.
Spoken-book sample vs a live booth read vs a Murf lesson
“AI audiobook sample” is a search, not a product. The decision is the file. The ElevenLabs column matches Text to Speech help and the live pricing page as of August 2026. A live booth and a Murf lesson stay qualitative here — we already wrote the course desk, and this page does not hop there.
| Criterion | ElevenLabs spoken-book sample | Record the author | Murf studio narrator |
|---|---|---|---|
| What the listener hears | A 60–90s first-page hook you can regenerate when the opening changes | The actual author, in the actual room, a file that ages when you rewrite page one | A directed studio narrator timed to slides — built for lessons, not a literary open |
| When it is the right buy | The listing needs a human-sounding first page and you will recut it | The sample already happens live and the words will not be rewritten | The file is a taught step on a deck, not a spoken book |
| What you re-do when page one changes | Edit the sentence, Generate Speech, replace the file in the sample slot | Re-book the author, the room, and an editor | Edit the block, regenerate, replace the audio on the deck |
| Tool this desk writes for | ElevenLabs Text to Speech — script, Voice Library, Generate Speech, download | A booth or a quiet room. Right when the voice has to be live | Ordinary path: the Murf course-lesson how-to. Not a hop on this page |
| Best 2026 fit | A retailer, Kickstarter, or author-site sample that should sound like a person reading a book | A one-time live read the author already records and will not rewrite | A 2–4 minute taught step — see the Murf course-lesson page |
ElevenLabs
A Text to Speech first page for a title that already has a sample slot. Free plan to audition the open; a paid plan is the commercial-rights download you can put on a store listing.
The script is a 60–90 second first page, spoken
Do not write a cold open for a documentary. Do not write a first-second Reel hook. Do not write a product close for a pricing page. Time the copy out loud. Sixty seconds is enough for the room and one name. Ninety seconds is one unfinished question. Longer than that and you are writing chapter one.
A working shape, spoken at a normal pace — roughly a minute and a quarter:
The harbor light on Wicklow Point had been dark for three nights when Mara found the envelope under the bait-shop door.
No stamp. Her name in pencil. Inside: a tide chart from nineteen eighty-seven and one sentence — the same sentence her father used to say when the fog sat on the water and he would not take the boat out.
Do not go past the salt line until the bell rings twice.
She looked at the chart. The salt line was not a metaphor. It was a depth mark, drawn in red, halfway to the wreck.
That is an audiobook sample. The first sentence puts the listener in a room. The middle names the person they will follow. The last sentence is an unfinished question. A trial URL on the end is a landing-page move — if you need a sales close, that is the sales-voiceover how-to, not this file.
Spell the words the model should say. “Nineteen eighty-seven,” not “1987.” “Wicklow Point,” the way you want those two stresses. ElevenLabs’ Text to Speech help is blunt about numbers and symbols: write them out, especially on multilingual models, because the same digit is pronounced differently across languages. A title called “Q3 ARR” needs those letters in the script the way you want them heard — and if your book is not called that, do not paste a product name into a literary open.
Keep one voice for this sample and for the recut. A library narrator that survives your place names is worth more than a cinematic whisper that flubs the last line. Emotional range is not why you are here. If realism-versus-timeline is the actual question, the comparison is ElevenLabs vs Murf vs Synthesys — ordinary path, no hop. Full product notes live in our ElevenLabs review.
Pick a voice. Generate the first page. Leave the last line alone.
Open Text to Speech. Select a voice from the control ElevenLabs documents at the bottom left — Default Voices or the Voice Library. Preview. Apply one voice and keep it. A listener who hears March’s sample and October’s recut should still recognize the person who said the place name.
Clone only if that person has to be you. Instant Voice Cloning is the published fast path from short samples (about 1–2 minutes of clean audio in the cloning help). Professional Voice Cloning is the dedicated model: Creator plan or above, a longer training set (published as 30–180 minutes), and a wait while it fine-tunes. We are not going to invent a clone wizard on this page. If listeners would notice a stranger and you do not already have the voice in My Voices, use How to clone your voice with ElevenLabs and come back.
Paste the first page as one short block. Voice, then model,
then settings — ElevenLabs ranks those in that order.
Multilingual v2 is the published stable default and the one
we would start on for a weekly English sample. Eleven v3 is
the expressive model: audio tags such as [sighs]
or [clears throat], a 5,000-character cap, and
fewer of the older sliders (Speed, Similarity, and Speaker
Boost are documented as unavailable on v3). Flash models are
the low-latency family; a seventy-second sample does not
need 75ms.
Where the sliders exist, the official starting point is Stability around 50, Similarity around 75, Style exaggeration at 0. Speed defaults to 1.0; the documented range is 0.7 to 1.2. Lower Stability for a livelier take, then generate more than once — the model is nondeterministic. Higher Stability for a straighter reader. Do not invent a slider we cannot see on the public page. If a control is missing on your model, change the sentence and re-roll.
Pacing is the sample-specific job. A dash or em-dash is the documented beat — use it before the last line so the hook does not run into the silence. Ellipsis adds hesitation; one is enough on a first page. Nudge Speed down if the opening room eats the name. On Multilingual v2, Flash v2, and Flash v2.5, ElevenLabs also documents an SSML break tag for a timed pause of up to three seconds. Confirm the live syntax in their Text to Speech help rather than pasting markup we cannot see unchanged. On Eleven v3, use audio tags and punctuation — that model’s help says it does not support SSML break tags.
Press Generate Speech. Listen on headphones, then on a phone speaker — that is how a store sample player sounds. Re-roll the first sentence until the room is clear. Leave a clean last line alone. Two free regenerations of the exact same text and settings are the published allowance; any edit to the copy or the sliders is a new generation.
Download the file, then put it where the sample already lives
The download is the gate. ElevenLabs’ Text to Speech help: after you generate, use the download control on the bottom right. Earlier takes sit in History on the same page — sidebar Text to Speech, then the history panel (or the history icon above Generate Speech on a narrow screen). History lists MP3 at 128 kbps or WAV; Advanced adds MP3 192 / 256 kbps, M4A, and FLAC. Higher-quality options are listed on paid tiers. For a retailer sample or a Kickstarter block, WAV or a high-bitrate MP3 is the usual handoff. Confirm the live list on the plan you pay for.
Text to Speech is an audio export. We do not invent an MP4 download on that playground. The listing is the slot you already have — the store player, the campaign audio block, the author-site embed — plus this file. Import the WAV. Leave a breath after the last line so the hook is not cut by the player’s fade. If you use a music bed under the sample, duck it so the first sentence is the loudest thing in the first second. We do not invent a store preset inside ElevenLabs. Confirm the live upload rules, duration caps, and format list on the retailer or campaign host you use.
Export is not a second ElevenLabs product. Name the source
audio so the next person can find it:
salt-line_sample_en_v1. When the opening
paragraph or the title page changes, open the sentence,
generate, bump to v2, replace the file. Leave the last line
that is still true. That is the reason you did not film the
author for the sample.
Listeners still deserve a plain-language note when the sample is synthetic. Say it in the retailer description or the Kickstarter story. If the current store composer asks you to label AI-generated or synthetic audio, use that control; check the live form, not a blog post from last year. That label is separate from the vendor license. You need both.
Commercial rights, the store listing, and the next cut
Read the live page. ElevenLabs meters in character credits across the suite. Free is 10,000 credits a month and no commercial license. Starter is the first paid tier that lists a commercial license and Instant Voice Cloning — $6/month on the public grid we verified in August 2026 (elevenlabs.io/pricing). Creator is the tier that lists Professional Voice Cloning ($22/month on that same grid; annual billing is cheaper). Confirm checkout. We do not print a commission rate.
A 60–90 second sample is a rounding error on credits. The plan question is the license, not the meter. Sit on a paid plan before the file is the one a store listing plays. If you later want the author’s likeness on the full title, that is the clone how-to, not a reason to start this brief in Creator. If the remaining job is a thirty-second product close, the sales-voiceover how-to is that brief. If the remaining job is a fifteen-second show open, the podcast-intro how-to is that brief. If the remaining job is a company-page post, the LinkedIn-voiceover how-to is that brief. If the remaining job is a Day-1 welcome, the onboarding-voiceover how-to is that brief. If the remaining job is a 9:16 hook, the Shorts-voiceover how-to is that brief. If the remaining job is a four-minute episode, the YouTube-voiceover how-to is that brief. If the remaining job is a taught step on slides, the Murf course-lesson page is that brief. The voice-tool comparison lives on ElevenLabs vs Murf vs Synthesys. This page does not hop there. The only Try button here is ElevenLabs.
Frequently Asked Questions
What is the best AI tool for an audiobook sample in 2026? +
Can I make an audiobook sample with ElevenLabs for free? +
Should I use a stock ElevenLabs voice or clone the author for the sample? +
How do I export an ElevenLabs audiobook sample as WAV or MP3? +
How do I tweak pacing on a 60–90 second first-page read? +
How is this different from the sales, podcast-intro, LinkedIn, onboarding, Shorts, and YouTube voiceover how-tos? +
Do I need to disclose that the audiobook sample is AI-generated? +
Continue the Pipeline
- Guide How to make a sales voiceover with ElevenLabs →
- Guide How to make a podcast intro with ElevenLabs →
- Guide How to make a LinkedIn voiceover with ElevenLabs →
- Guide How to make an onboarding voiceover with ElevenLabs →
- Guide How to make a YouTube Shorts voiceover with ElevenLabs →
- Guide How to make a YouTube voiceover with ElevenLabs →
- Guide How to clone your voice with ElevenLabs →
- Review ElevenLabs voice cloning review →