There’s a durable little stereotype that an audiobook is a print book wearing headphones.
The same object, only now it has been asked to speak clearly and not breathe too loudly.
Which is how a perfectly respectable book ends up cast as its own understudy.
An audio edition is a new production.
It uses the same script, certainly.
But it needs a right to mount the play.
A territory in which to sell tickets.
A payment structure that survives rehearsal.
A delivery route.
And a performer whose first minutes make the audience believe the promised part has actually entered the room.
The trade collapses all of that into the phrase doing audio, which has the administrative efficiency of calling a production of King Lear some weather and a family disagreement.
The useful distinction is sharper.
A grant permits exploitation. A performance creates format value.
One without the other gives you either a splendid audition for a part nobody owns, or a very expensive microphone pointed at a contract.
So the ten-second question is: who holds audio rights here?
The five-minute question is whether the voice makes your jacket promise audible.
They belong together.
Here is how to cast the book.
Read the programme first
Before anyone starts discussing voices, inspect the programme.
A grant-of-rights clause moves a defined slice of your author’s bundle to a publisher, for a defined territory, format, and term.
It is not a mood board.
A sound producer cannot cure an absent grant with an excellent Welsh accent.
A well-drafted clause separates print, ebook, and audio.
It also names what stays reserved.
A right you do not name is a right you have not granted.
Territory needs its own line of sight. US and Canada, the UK, and World English are familiar shapes, but they’re different stages. Not alternative spellings of wherever listeners happen to be.
Language is another cue to read rather than infer.
So is term. The usual full copyright term runs for the author’s life plus seventy years, which is a very long run for a show nobody has yet cast.
None of which demands that every audio deal become a small constitutional convention.
It insists that format, territory, language, and term answer the same practical question.
Who may put which performance in front of which listener?
If your domestic publisher holds US and Canadian audio but not World English, a production that sounds built for a wider English-language route still needs the rights chain to reach that audience.
The stereotype says the audio right is a file attached to the hardback deal with a paper clip.
It is a format grant with a distinct commercial life.
A publisher holding it may produce the edition directly.
Or license it to an audio specialist.
A retained right may be produced by the author or a small press through a direct-to-platform route.
Neither is intrinsically the grander casting decision.
The relevant question is which arrangement can actually put the right performance in front of the right market.
Do the first five minutes enter on cue?
Casting begins months before a studio door opens.
You choose single narrator or full cast, then audition against a few manuscript pages.
Dialogue. An accent. A difficult name.
The passages most likely to expose whether a performer can keep the part upright.
The big audio houses now run talent platforms specifically to widen the group considered for those auditions.
Where author involvement is negotiated, approval of the tape often comes before signature, because a narrator becomes attached to a listener’s memory in a way the person who adjusted the kerning mercifully does not.
The first-five-minute contract is not can this person pronounce the book?
A text can be pronounced and still arrive in the wrong costume.
Consider two short personal essays about a daughter clearing her father’s flat.
The first is recorded by a human narrator whose restraint makes each object feel like a decision not yet explained. The silences are part of the family argument.
Your listener receives intimacy, pressure, and a reason to stay with a difficult consciousness.
The jacket promise has entered on cue.
Now give the same essay a synthetic read that makes every sentence equally settled, each revelation carrying the tonal distinction of a supermarket announcement.
The text has not become worse.
The production may be perfectly intelligible.
But your listener route has changed, because the edition no longer delivers the particular proximity its presentation promised.
That is not a moral verdict on a technology, or a claim that one kind of listener is more cultivated than another.
It’s format value. The performance altered what the edition is selling.
And the evidence makes this less theoretical than a panel discussion with excellent lanyards.
Synthetic narration is a rounding error in audiobook revenue.
Roughly a sixth of listeners have tried one, and stated willingness to try has been falling rather than rising.
The interesting gap is not whether synthetic narration exists.
It’s whether a given performance carries the listener promise well enough to earn a repeat booking.
So make the test deliberately local.
For an intimate essay, a book whose voice is among its attractions, or a dialogue-driven novel: does five minutes justify the jacket?
For a practical backlist edition whose promise is access to information, another performance may clear that same test easily.
The answer positions the edition.
It does not put the manuscript on a moral ladder.
Price the rehearsal, not just opening night
The money has to agree with your casting.
Narrator pay is quoted per finished hour, not per hour inside the booth.
Which matters, because production accounts put the work at two or more studio hours for every finished hour, before pre-production.
A ten-hour finished novel is not a pleasantly long afternoon with a glass of water and some enthusiastic consonants.
The available ranges are wide enough to be an acquisition fact.
Budget voice actors run a few hundred dollars per finished hour.
The union minimum sits near the bottom of that band, plus an employer contribution to health and retirement.
Brand-name narrators command a thousand or more, roughly five or six times the union floor.
So your ten-hour novel runs from around fifteen hundred dollars at the floor to five figures at celebrity rates, before editing, proofing, and mastering.
Audio is cheaper than print is not an analysis.
It’s a shrug in a headset.
And those last three tasks are not a single post-production blur.
Editing removes breaths, mouth noise, and flawed takes.
Proofing follows the manuscript line by line, and can send a missed word back to a pickup session.
Mastering normalises levels, sets the noise floor, and converts the result to a platform’s specification.
Studios price all three separately, which is a useful reminder that the jobs stay separate even when a budget has wished very hard otherwise.
Full cast changes the rehearsal calendar again. Every speaking role needs time, and sessions may need sequencing or layering.
Which is why it costs more and takes longer than a single narrator at comparable runtime.
The question isn’t whether a chorus of talented people beats one performer.
It’s whether your book’s listener promise is made of exchanges that need a company, or of a consciousness that needs one person at the lectern.
Synthetic production is a different pricing structure, not merely a smaller human invoice.
Some services charge per finished hour.
Others offer whole-book narration at a flat price regardless of length, across dozens of languages.
A flat per-book price favours a long book in a way per-finished-hour never does.
That can make an edition commercially possible.
It also makes the five-minute test more important, not less, because the saving and the appeal are being generated by different systems.
Who carries the sales risk offstage?
Payment structure is part of casting, because it declares who is being asked to believe in the run.
Flat per-finished-hour payment is work for hire. Your narrator is paid on delivery, and you hold the sales risk.
Royalty share moves much of that risk onto the narrator, in exchange for a claim on long-tail revenue.
It is not a kindly discount for the person doing the work.
It is a bet.
Under the common exclusive royalty-share structure, the total royalty divides evenly between author and narrator.
That can be entirely rational for a narrator assessing a backlist author with a track record, a strong platform, or a genre with dependable audio attachment.
For an unproven title that sells modestly, it can pay less than the union floor.
Experienced narrators aren’t being difficult when they decline such offers.
They’re reading the box-office forecast on your playbill.
The same distinction belongs in your rights negotiation.
The widely used model contract splits audio-recording income evenly between author and publisher.
Self-production may leave your author with all of a smaller, self-funded stream.
Neither percentage is a magic word.
The comparison is a share of a producer’s reach versus ownership of a producer’s risk.
A publisher with real audio distribution can make its half earn more in absolute dollars.
A holder with no plausible route to listeners is merely reserving a theatre.
And subsidiary income doesn’t stroll in wearing a separate envelope.
For a right the publisher controls, it returns through the original publisher, divides under the contract, and credits the same royalty account.
If the advance hasn’t earned out, it reduces the unrecouped balance first.
Which means a late audio sale can be the event that earns out a title whose print performance had gone quiet.
And why an unused audio grant stays valuable to its holder long before anyone has booked a narrator.
Check the exit route
The finished master still needs a theatre.
A platform’s delivery specification can require a particular format, sample rate, noise-floor and loudness tolerance, and chapter-marker structure.
A failed file comes back to you for another mastering pass.
Some platforms have historically required human narration and rejected synthetic output at exactly that technical gate, which makes it a current-policy check before any production decision.
Not a historical curiosity to frame above the espresso machine.
Then comes the schedule, which is the decision most often made by default.
Audio can launch with print and ebook.
It can arrive later, to protect hardcover sales.
It can precede print.
It can be the only primary edition.
Audio-first revenue jumped by half in a single recent year, reaching six per cent of total audiobook net revenue.
That is not an ebook with a better sound system.
It’s a show with no print opening night to wait for.
And the broader market gives the decision its scale and its caution.
US audiobook sales are in the billions and still growing at high single digits.
The number of active titles grew by more than forty per cent in a single year.
Well over half of American adults have listened at some point.
The average listener finishes fewer than four titles a year.
The audience is real.
So is the competition for one evening’s attention.
A broad market is not an automatic standing ovation.
Check the rights after the curtain falls
The inconvenient case is the book with a live ebook, some print-on-demand availability, and no meaningful audio activity.
It looks in print from the lobby.
It may be commercially silent in the auditorium.
A reversion clause exists because those are not the same condition, and its detail belongs in its own conversation.
What matters here is the distinction between a desired new production and an available right.
Your author might hear the perfect voice for a dormant essay collection. Or notice that a previous audio version never carried the book’s promise.
If audio was retained, you can consider a fresh licence directly.
If it was granted, the work begins with the contract and its reversion machinery.
Not with an audition tape emailed at 1:14 a.m. to a narrator who has unfortunately become the inner voice of the project.
And the ambiguous case is not an exception to the instrument.
When a human performance makes a book intimate and a synthetic version makes it legible but changes its appeal, the synthetic edition may still have a route.
Treat the change as an adjacent edition, with its own cost, listener promise, platform eligibility, territory, language, and rights analysis.
It is not a counterfeit print book.
It is also not the same product wearing a different headset.
Run the cast list
Several of these should be true before an audio migration begins:
- The grant explicitly names audio, and its territory, language, format, term, and reservations match the production you’re planning.
- A five-minute tape from the proposed narrator or cast makes the edition’s jacket promise audible.
- Your budget distinguishes per-finished-hour pay from studio time, editing, proofing, mastering, pickups, and any full-cast scheduling cost.
- A flat-fee or royalty-share offer says plainly who bears sales risk and what the long-tail split actually is.
- You have chosen your delivery route before mastering to a platform’s technical gate.
- The release is scheduled as simultaneous, delayed, backlist, or audio-first, rather than treated as a file that simply appears after publication.
- Retained rights, royalty routing, audit exposure, and the reversion test have all been read before a dormant right gets called active.
The ideal result is not that every book gets audio.
That would give you a very crowded repertory company and some deeply regrettable casting.
It is a clear answer about market position.
Does this book have a listener route the rights can lawfully reach, the economics can finance, and the performance can make desirable?
Audio migrates a book into a new market only when those answers agree.
The contract has to know who owns the part.
The performance has to know who speaks it.