Skip to content

Voice-Over Production: A Brand's Guide to Getting It Right

By Jason Kidd 12 min read
Voice-Over Production: A Brand's Guide to Getting It Right: Killerspots branded jingle & audio production graphic

I have been casting voice talent for something close to three decades, and the same failure repeats at the front of most projects. A brand books a session, hires a genuinely good actor, gets back a clean take, and the finished spot still does not land. Nobody made a mistake in the room. The mistake happened before anyone reached the room.

Voice work looks simple from the outside, because the deliverable is one person talking. What you are actually buying is a chain. A script written to be spoken, a casting decision that matches what the listener already expects, direction that produces a performance instead of a recitation, and a space clean enough that the take survives the mix. Break any link and you pay for it twice, once in pickups and once in a spot that never quite works. Here is how voice-over production really runs, and where brands most often lose the thread.

What does voice-over production actually involve?

Direct answerVoice-over production covers everything around the recording, not only the recording. It means a script built for the ear, casting matched to the audience, direction aimed at a performance, capture in a treated room, then editing, mixing, and delivery to whatever technical spec the destination requires.

The recording itself is the shortest part of the process. A 30 second read with locked copy and the right actor can be finished in under half an hour. Everything that determines whether it works sits on either side of that half hour.

On the front end: copy that can be performed, a clear picture of who is listening, and a casting choice. On the back end: editing out breaths and mouth noise, mixing against music and effects, and delivering a file that meets the loudness and format requirements of wherever it runs. Broadcast has real technical standards, and a mix that sounds great on studio monitors can be rejected or quietly turned down by a station’s processing if it ignores them.

Brands tend to shop for the middle piece, the booth time, because it is the visible part. The parts that decide the outcome are the ones nobody photographs.

How do you write a script a voice actor can actually perform?

Direct answerWrite for the ear, not the page. Short sentences, one idea each, contractions, and no clause a listener has to hold in memory. Read it aloud on a stopwatch before the session. If you stumble on your own copy, a professional will stumble too, and fixing it costs session time.

The most common script problem is that it was written to be read silently. Written English tolerates long subordinate clauses and stacked qualifiers. Spoken English does not, because a listener cannot go back a line. If a sentence needs a comma to survive, it probably needs to be two sentences.

The specific things that cost time in the booth are predictable, and you can catch every one of them by reading the copy out loud:

  • Sibilance stacks. Several s sounds in a row (“seasons of savings starts Saturday”) whistle on almost every microphone and take real work to tame.
  • Tongue twisters nobody noticed. They are invisible on screen and obvious the moment somebody speaks them.
  • Numbers, URLs, and legal copy. These consume far more time than their word count implies, because they have to be delivered slowly enough to be understood on one pass.
  • Two ideas in one sentence. The actor has to choose which one to stress, and whichever they pick, half the room will disagree.

Timing matters more than writers expect. A comfortable 30 second read runs about 65 to 75 words. Copy that arrives at 95 words is not a 30 second script yet, and the decision about what to cut should be made by you at a desk, not by an engineer with the clock running.

How do you cast the right voice?

Direct answerCast the listener's assumption rather than your own preference. The real question is who your audience expects to hear and believes immediately, which is often not the voice the business owner likes best. Audition against your actual script, because a demo reel proves range while a custom audition proves fit.

Demo reels are useful for building a shortlist and almost useless for making the final call. A reel is an actor’s best moments from their best sessions, assembled to show range. It tells you what somebody is capable of. It does not tell you how they sound reading your particular copy about your particular business.

Custom auditions solve this. Send three or four finalists the real script, listen to what comes back, and the decision usually makes itself. What you hear is not just tone. It is whether the actor understood the copy, which is the single best predictor of how the session will go.

The trap to watch for is owner taste. A business owner naturally gravitates toward the voice they would want to be, which skews toward authority and polish. Audiences frequently trust something closer to a competent neighbor. This is the same shift behind brands moving away from the classic announcer voice, and it is worth resisting the instinct to cast the biggest voice in the pile. Casting decisions should be made against the audience, not the org chart.

We keep a working roster rather than a list of favorites, and we cast against the brief every time. That matters for consistency too. If the same voice is going to carry your radio commercial, your television commercial, and your on-hold messaging, you are casting a long relationship, not a single read.

What makes a direction session work?

Direct answerDirection that names an intention instead of a sound. Telling an actor to add energy produces louder, not better. Tell them who they are speaking to and why it matters, give one adjustment per take, and record two alternate reads while the actor is warm so you have options in the edit.

“More energy” is the note that ruins more sessions than any other. It describes a symptom, and the only thing an actor can reliably do with it is get louder and faster, which is rarely the actual problem. What people usually mean is that the read sounds like an advertisement instead of a person.

Better notes are situational. Tell the actor they are explaining this to a friend who is skeptical. Tell them the listener is in a truck at seven in the morning and has heard four other ads already. Give them a person and a moment, and the performance adjusts on its own.

Two practical rules save real time. Give one note per take, because stacking three adjustments means the next read will address one of them and lose the other two. And record alternates while the actor is warm: a straighter version, a warmer version, one with a different emphasis on the tag. Those alternates cost a few minutes on the day and repeatedly save a second session later when somebody in approvals wants to hear a different feel.

Finally, decide who is directing before the session starts. Direction by committee produces a compromise read, and compromise reads sound exactly like what they are.

Why does the room matter more than the microphone?

Direct answerBecause a mediocre microphone can be improved in the mix and a room cannot be removed. Reflections, air handling, and street noise bake permanently into the file. A modest mic in a treated space beats an expensive one in a spare bedroom every time, which is why the space is the first question to ask.

Home studios became normal in this business several years ago, and plenty of them are excellent. Plenty of others are a good microphone in an untreated room, which is the audio equivalent of a great camera pointed at bad lighting.

What a bad room does is add a signature you cannot subtract. Hard parallel walls produce a boxy quality. Too much soft furnishing produces a dead, muffled read. A refrigerator or an HVAC cycle raises the noise floor, and once it is in the recording underneath the voice, removing it removes part of the voice with it. Modern processing tools are impressive and still not magic.

There is a consistency argument too, and it is the one that bites brands later. If you record a campaign now and need a pickup line in four months, that new line has to sit invisibly next to the original. Matching a read across time requires the same voice, the same room, the same signal chain, and the same engineering approach. This is a large part of why we track in our own rooms and keep the session files: a matched pickup six months out is routine when the setup is documented and difficult when it is not.

What should you settle about usage rights before the session?

Direct answerSettle term, territory, and media in writing before anyone records. A read licensed for regional radio does not automatically cover television, streaming, or a national rollout. Brands get surprised when a campaign succeeds and expands, because expansion is precisely what the original license did not include.

This is the part brands skip and later regret. Voice talent is not licensed like stock music. The fee attaches to how the recording is used, and three variables define it: how long you may run it (term), where it may run (territory), and in what media (radio, television, streaming, digital, internal, industrial).

The failure pattern is always the same, and it is a success story going wrong. A spot performs, so the client wants to extend the flight, add three markets, or cut the audio into a pre-roll video. Every one of those is a different use than the one that was licensed, and renegotiating from a position of “we are already running it” is the weakest possible spot to be in.

Two habits prevent all of it. First, tell whoever is handling casting what you might do with the audio, not only what you have committed to today. Buying wider usage upfront is nearly always simpler than expanding a license mid-flight. Second, keep the license terms filed with the audio itself. Agencies change, staff turn over, and in two years the only thing that will settle a question is the paperwork.

Should you use an AI voice or human talent?

Direct answerUse synthetic voice where the copy is functional and disposable, such as internal prompts, prototypes, or high-volume variants. Use human talent anywhere the read carries persuasion or brand trust. Synthetic voices handle information well and intention poorly, and intention is the entire job in a commercial.

I will give the honest version, including the part that does not favor my own department. Synthetic voice has become genuinely good at delivering information clearly. For a scratch track to time a script, an internal training module, an interface prompt, or a set of near identical variants across many product names, it is a reasonable tool and nobody in the audience is being asked to feel anything.

Where it still falls down is subtext. Commercial reads work through implication: the small hesitation before the important word, the warmth that makes a claim sound like a person’s opinion instead of a corporate statement, the restraint that keeps a line from sounding like a pitch. Those choices come from an actor understanding why the line exists. That is also what makes voiceovers effective in advertising in the first place, and it is not a technical gap that better rendering closes.

There is a rights dimension worth naming as well. Cloned voices raise consent and licensing questions that are still being worked out, and a brand does not want to discover the boundary the expensive way. We cast professional human talent for client-facing work for both reasons, the performance and the paperwork.

When is a professional voice-over the wrong call?

Direct answerWhen the piece does not need persuasion. Internal training, temporary scratch tracks, and quick social cuts where a real employee on camera reads more credibly than a polished announcer. Hiring a pro for those spends budget on polish the audience was not asking for and can actively hurt authenticity.

The clearest case is owner and staff content. If a customer is watching a short video of the person who will actually show up at their house, a broadcast-quality announcer voice over the top makes it feel produced, and produced is the opposite of what that format is selling. Let the real person talk.

The same applies to fast social content with a short shelf life and to anything you are still testing. Do not commission a finished read for a script that has not proven it works. Rough it in, run it, and invest in the performance once you know what the message is.

Where the calculation flips is the moment you are buying media. If money is going behind the impression, the read is the impression, and the difference between a professional performance and a passable one shows up directly in response. The same logic applies to a custom jingle or any other asset that will run for years: the production cost happens once and the exposure keeps compounding.

How do you set up your next session to go well?

Direct answerLock the script and time it out loud first. Define the listener in one sentence. Audition finalists against the real copy. Name one decision maker for the session. Settle usage in writing. Record alternates while the actor is warm. Those six steps prevent nearly every avoidable pickup.

None of it is complicated, and all of it is easier before the session than after. Voice-over rewards preparation far out of proportion to the effort it takes. An hour spent reading copy aloud and deciding who the listener is will save you a second session and produce a better spot at the same time.

If you would rather hand the whole chain to somebody who does it daily, that is what our voice-over production work covers: casting from a working professional roster, direction, recording in our own rooms, mixing to broadcast spec, and usage handled correctly upfront. It sits alongside the rest of our audio production capability, so the voice on your commercials, your jingle, and your phone system can be the same voice telling a consistent story. Tell us the script and where it needs to run, and we will handle the rest.

Frequently asked questions

How much does voice-over production cost?

There is no flat rate, because the fee is driven by usage rather than by studio time. The variables are how long the spot runs, where it airs, which media it covers (radio, television, streaming, internal, web), the length of the license term, the talent you cast, and how fast you need it. A regional radio read and a national television campaign using the identical script and the identical actor are priced very differently, and that difference is the license, not the recording. Send us the script and where you intend to run it and we will quote the whole package, session and usage together, so nothing surfaces later.

How many words is a 30 second voice-over?

Roughly 65 to 75 words for a comfortable conversational pace, and about 80 to 85 if the read is energetic and the copy is simple. Those numbers are a starting point, not a rule. Phone numbers, web addresses, legal disclaimers, and product names all eat more time than their word count suggests, because they have to be delivered slowly enough to be understood. Always read the copy out loud on a stopwatch before the session. If you are cutting words in the booth, you are paying studio time to do work that belonged in the draft.

How long does a voice-over session take?

A single 30 or 60 second spot with a finished script usually takes 15 to 30 minutes of recording, including alternate reads. Longer form work such as narration, e-learning, or a full on-hold program runs closer to an hour per finished ten minutes, because there is more pickup and more breath editing. The part that expands unpredictably is approval. Sessions run long when the script is still being rewritten in the room or when several stakeholders are giving conflicting notes at once, which is why we lock copy and name one decision maker before the actor opens the mic.

Do I need to be at the voice-over session?

It helps, and you do not have to be in the building. Most brand sessions run with the client patched in remotely, listening live and giving notes between takes through the engineer. That gets you the read you actually wanted on the day rather than a revision cycle a week later. If you cannot attend, send a reference: a spot whose tone you like, or a note describing the listener and the moment. Directing from a written brief works fine when the brief is specific about intention, and poorly when it only describes a sound.

Can I use the same voice-over in more than one ad?

Only if the license says so. Reusing a read in a different medium, a different market, or after the term expires is the most common way brands end up in an uncomfortable conversation, and it is entirely avoidable. If you expect to cut the audio into several versions, run it in more than one market, or keep it running past a season, say so before the session so the usage gets written correctly the first time. Buying the wider license upfront is nearly always simpler than renegotiating a license for a campaign that is already working.

Want results like this for your brand?

Killerspots is a full-service creative + digital agency. Let's talk.

Get a Free Quote
LeadConnector

Capture every lead. Follow up automatically.

LeadConnector — our AI-powered CRM — captures the leads your marketing drives, scores them by intent, and follows up 24/7 by text and email. Missed call? It auto-texts back. No lead ever goes cold.

AI Lead Scoring 24/7 AI Follow-Up SMS + Email Unified Inbox