What Happens to Your Audiobook Between the Booth and Your Inbox

Inside the audiobook recording booth at Toe-Curling Tales.

(Yes, I cleared my desk before the photo. The piles of loose-leaf paper and half-empty notebooks are just temporarily on the floor, until I get tired of stepping around them.)

A look at the seven phases of production behind every finished hour of audio at Toe-Curling Tales, and where you come in at the end.

My audiobook narration FAQ says approximately six hours of work for every finished hour of audio. Let’s walk through what those six hours actually mean.

Maybe you’re considering hiring me. Maybe you already have and want to understand what’s happening on my side. Either way, this is the full picture. Each phase exists for a specific reason, and understanding those reasons makes the whole process easier to navigate, especially when it’s time for your own review at the end.

Phase 1: Prep Read

Before I record a word, I read your full manuscript—and I do mean a full read, the way I’d lose myself in a book off the shelf. This is where I get a feel for your characters’ attitudes and communication styles. I internalize their emotional arcs and rapport with each other. I work out the physical situation each scene puts them in, because that rewires a performance before the emotion even gets a vote: a character crawling under a flipped car to reach someone trapped inside doesn’t sound like the same character having a heated conversation while standing up—the breath, the strain, and the volume all shift. And I research pronunciations, flagging every word I’ll need to confirm with you before recording.

For speculative fiction, the flag list gets long: “uncommon” names (both proper names and the worldbuilding terms specific to your world), phonetic spellings that may not map to “standard” English phonics, and incantations or spell words. All of it needs to be confirmed before I step into the booth. This is why the pronunciation guide matters so much, and why authors who provide a thorough one protect their own story and investment. A narrator guessing at a name is a narrator who introduces errors the proofer has to catch later, stretching the production timeline. (No one has time for that.) A narrator with a clear guide ideally gets it right the first time.

My book editing background sharpens this prep read. I notice continuity threads, dialogue and attitude patterns, and (the dreaded word) subtext that shape how I’ll perform your book.

Phase 2: Recording

I record continuously and use a finger snap to mark retake points. When I catch a mistake mid-read, I snap my fingers (or clap because my fingers are tired) to mark my place in the recording, start again from a clean point, and continue. (Some narrators use a dog clicker for the same purpose; I use my hands because I tend to act with my hands, and, personally, holding a clicker distracts me.) This way, marking a retake doesn’t pull me out of the emotional state of the scene, and I know exactly where to pick up again as cleanly as if I’d done punch-and-roll (where you stop, back up, and re-record over the mistake in real time). I admire the narrators who can; it’s just not my method. The listener never hears the retakes. When I get to editing, the snap tells me exactly where to cut, and the corrected take is already there waiting.

Narration is cognitively expensive. I’m reading ahead, performing the current line, keeping each character’s voice consistent, managing my breath, and landing the right pronunciation and emotion. Even experienced narrators produce several misreads per finished hour because sometimes the brain auto-completes a grammatically plausible word that isn’t the one written. And sometimes I’m just too far inside the scene—too excited, too angry, and too heartbroken to land clean, because the performance is pulling on something real in me, not just on the page. This is the reality of the work, and the production pipeline exists specifically to catch and correct these.

For scale: a ten-hour audiobook might take thirty to forty hours of booth time. “Just read it out loud” is not what narration is. (Trust me, I once fell for the “get rich using your voice” gimmick too.)

Phase 3: Narrator Self-Proof

I do my own proofing pass by listening to the recorded audio against the manuscript to find as many of my own errors as I can: dropped words, substitutions, and mispronunciations I didn’t catch in the moment.

By the time the files leave my hands for the proofer, your audio has already been through one quality pass. The raw recording is messier than anyone downstream ever hears. (Especially since I don’t do punch-and-roll: a thirty-minute chapter can be an hour and a half of me getting there.)

Here’s the honest part: I can’t catch everything in my own work. By the time I’m self-proofing, I’ve read your book twice, and now I’m partway through a third. My brain knows your text cold, and a brain that knows the text fills in what it expects to hear. That’s not carelessness; it’s how attention works when you’re inside the material. Phase 5 exists because an outside ear doesn’t have that disadvantage.

Phase 4: Technical Editing

Full transparency upfront: Engineering is a different craft from narration. There are MANY who are talented and can utilize both sides of their brains to successfully do both. I am a narrator, which is exactly why I invested in working with a professional audio engineer to build a custom processing stack tuned specifically to my voice and my recording environment: EQ (boosting or trimming specific frequencies so the voice sounds clean and natural), compression (evening out the whispers and shouts, aka quiet and loud sections so you don’t have to keep playing with the volume button to listen), and limiting (a hard ceiling on volume so the loudest moments can’t spike into distortion). That way, the technical side of the work serves the performance rather than fighting it. I then run every audiobook through that stack and handle all the editing and mastering myself.

The cleanup pass handles mouth noise and clicks (have you ever noticed the wet squelch or saliva bubbles popping when you open your mouth?), plosives (the popped P’s and B’s), sibilance (harsh S sounds), and breath spikes (I ran out of air). I manually match room tone across chapters recorded on different days. I tighten the gaps between sentences without flattening the natural rhythm of the performance. Some editors work to a rule of thumb: the pause between sentences runs about the length of a natural breath, half a second to a second, adjusted for tone. That’s not wrong; it’s just not mine. I trust that I was in the moment when I recorded the scene, and that I paused as long as I did for a reason, emotional or technical. My job is to protect that pause, not standardize it.

This cleanup pass also includes checking that every retake splice gets smoothed. When I correct a mistake mid-scene, there’s a seam between the original take and the corrected one. It’s my job to make every seam inaudible.

Mastering brings the audio up to retail spec: the loudness, balance, and consistency that let your book play cleanly across headphones, car speakers, and phone speakers alike.

After the editing and mastering work is done, I do a technical verification pass before anything moves on to the proofer. While the custom processing stack is consistent, “consistent” isn’t the same as right for every passage. Sometimes a loud emotional line gets over-compressed. Sometimes the stack cuts a breath when it was intentional. My ear is the safety net for my own technical work, just as the proofer’s ear is the safety net for the manuscript-to-audio match. Two QC passes with two different jobs.

Phase 5: Proofing

A dedicated proofer listens to the edited and mastered audio while following the manuscript, hunting for any textual errors that survived my self-proof.

What they’re catching: dropped words, substituted words (“will” when the manuscript says “while”), added words, skipped lines, mispronunciations, singular-plural mismatches, wrong verb tenses, and anything where the audio doesn’t match the text.

My go-to proofer also flags performance notes, not just textual errors. They’re the only person in the production chain who experiences my performance the way a listener would, from start to finish in real time, without already knowing what’s coming. They’re not reviewing your book; the book is the reference. They’re reviewing my take on your book, which means they also catch what a pure text check wouldn’t: moments where my delivery didn’t serve the scene the way the chapter set it up to, or character voice that drifted from the established direction. Those notes don’t get treated like misreads. I take them as input, listen back to the section, and decide whether a pickup serves the listener better than the original take. Sometimes I agree and re-record, and sometimes I trust the original choice and we leave it as is.

Phase 5 protects your words. Every error that slips past two passes of my own ears meets a fresh set of ears hearing the book for the first time. It’s the most labor-intensive step in the chain, and the one that matters most for textual accuracy.

One thing worth knowing: ACX’s own quality review checks technical specs (volume levels, noise floor, file format) but does not perform manuscript-to-audio proofing. An audiobook can pass ACX QC even if it contains dozens of misreads. The textual accuracy of your audiobook falls entirely on the narrator and their proofer. This is where mine gets defended.

Some proofers use AI-assisted tools (like Pozotron) that compare the recording against the manuscript and flag where the two diverge. To be clear, this is proofing software, not an AI voice. I perform every word myself, and a human verifies every flag it raises. My default is a human proofer. AI-assisted proofing is available only if you ask. Either way, the output is a corrections log: a document listing every error with chapter, timestamp, what the manuscript says, and what I actually said.

Phase 6: Pickup Recording and Final Mastering

I receive the corrections log and then re-record every flagged error. These are called pickups, and each one needs to match the original performance in tone, energy, pacing, and room sound, even if weeks have passed since the original session.

I listen back to the original audio around the correction, get into that exact vocal placement and emotional register, record the corrected line, and splice it in seamlessly. Some pickups take thirty seconds; some take ten minutes of takes to get the match right. It’s why “can you just fix that one word?” is never as quick as it sounds, because every correction means rebuilding the performance around it.

Once pickups are recorded and spliced in, I remaster the affected files to meet retailer technical specifications, and then run another technical verification pass to confirm the corrected audio matches the spec of everything around it. For ACX/Audible, that means RMS between -23dB and -18dB, peaks below -3dB, a noise floor below -60dB, plus chapter head and tail silence, file format, and naming conventions. ACX is usually the standard, but let me know if you want to upload to a specific retailer.

Phase 7: Author QC (Where You Come In)

You have 10 business days to listen to the mastered files and flag any remaining issues.

What to listen for:

Anything that breaks the spell: a mispronunciation, a technical issue, or a moment that doesn’t land the way you wrote it. You know this book better than anyone alive.

What author QC is not:

A second pass at direction. The artistic approach (tone, pacing, and characterization) was approved during the fifteen-minute sample phase (or, if you purchased the Premium Checkpoint Review Service, throughout production). Author QC is for catching genuine problems, not for rewriting performance choices you already signed off on.

If you find a mistake, please fill out the QC Corrections Log with the following: chapter #, timestamp, text as written, what was said, and why it needs fixing.

Audiobook QC corrections log template—spreadsheet for flagging misreads, mispronunciations, and audio issues. Columns: chapter, timestamp, manuscript text, audio text, notes.

Audiobook QC corrections log template—a spreadsheet for flagging misreads, mispronunciations, and audio issues, with columns for chapter, timestamp, manuscript text, audio text, and notes.

Here’s why specificity matters: by the time your book reaches you for QC, I’ve read it, narrated it, and proofed it, and the proofer has gone through it word by word. We’ve lived in your book for weeks at this point. A note like “something sounds off in chapter 12” requires someone to re-listen to the entire chapter and hunt for the problem. A note like “Chapter 12, 14:55 said ‘gripped’ but manuscript says ‘grabbed’” takes thirty seconds to address. The cleaner your log, the faster your corrections get done and the less time everyone spends hunting.

What You Can Do to Make the Process Smoother

These come directly from what I’ve learned in production:

Finalize your manuscript before the recording starts.

Changes after the recording starts cost time and money. I’m not being rigid. Every change requires a re-record, a re-edit, and a remaster of the affected section. A manuscript that’s truly record-ready is the single biggest thing you can hand me.

Provide a thorough pronunciation guide.

Every invented word, character name, place name, and unusual term with phonetic spelling or (highly recommend) a short audio clip of you saying it.

Trust the production timeline.

Fourteen to seventeen weeks for a ten-hour book is the real math of a multi-pass pipeline, not padding. Quality takes the time it takes.

Understand that narrator errors are normal, not negligent.

Several misreads per finished hour are the cognitive reality of the work, not a measure of care. The production pipeline exists specifically to catch and correct these before they reach you. That’s what you’re paying for.

Why I’m Showing You All of This

My brand promise is that it’s your story and I’m the vessel bringing it to listeners’ ears. This is how I keep that promise.

The prep read exists so I know your world before I live in it. The recording approach exists so your listeners hear a performance, not a struggle. The self-proof, technical editing, proofing pass, and pickup cycle exist so that what listeners hear is exactly what you wrote. The mastering exists so your book sounds the way a retail audiobook should sound. Author QC exists so you get the final word on your own story.

My job is interpretation, performance, and the technical craft that makes them sound effortless. The proofer’s job is to make sure no word slips between manuscript and audio. Your job is to provide clear source material and honest, specific feedback during QC. When all three roles hold their own weight, the result is an audiobook that sounds like your book was always meant to be heard out loud.

That’s what I’m working toward, every time.

— Ashley

Previous
Previous

Every Word Counts · A free tracker for writing groups