Record a Scratch Voice Track That Lets You Judge the Commercial's Timing
Record a Scratch Voice Track That Lets You Judge the Commercial’s Timing
A scratch voice track exists to make one argument possible: whether the copy fits the picture. It is not an audition, not a placeholder to be defended, and not the first pass at the finished voice. You record it so that you, a director, or a client can hear where a line lands against a cut — and so that when the answer is "not quite," there is a file that shows exactly which words changed.
The awkward part is that a scratch take has to be two contradictory things at once: clean enough to trust, provisional enough to throw away. Get it wrong in one direction and nobody can hear the phrasing through the room sound, the puffs of air, and the clipping. Get it wrong in the other and someone starts treating the file as the voice of the commercial — and now you are arguing about a performance that was never meant to be one. The test is narrow and useful: nobody in the review should have to ask whether they are listening to the final read.
("Scratch" comes from music and radio — a working pass that exists in order to be replaced.)
Two things this recording does not do. It does not establish what a professional performer must deliver, and its duration does not fix how long the finished spot will be. A scratch read is one person, on one afternoon, under one set of circumstances. It answers "does this fit the cut in front of us?" and nothing beyond it. Most of the decisions below follow from keeping that boundary intact.
Decide which timing question the take has to answer
Three different questions get called "does the voice fit," and they need different recordings.
The first is total length: does the whole passage overrun the window the picture leaves for it? The second is landing: does one particular line arrive with the shot it is written against, or does it trail into the next one? The third is comprehension: does the phrasing you intend — where the stress falls, where the pause sits — actually make the sentence mean what you meant?
You can settle the first question with a stopwatch and a text-to-speech read, and honestly you should, because it costs nothing. Record yourself when the second or third question is live, because those are about where the emphasis lands, and no synthesizer will tell you that.
Before you roll, separate the copy into fixed and provisional. A product name, a legal phrase, a line the client already approved — those are fixed, and a scratch take cannot change them. The rest is provisional, which is the whole reason you are recording. Write the split down, because the moment a take is too long, everyone in the room starts trimming whatever is nearest, and the approved line is often nearest.
Choose the input you mean to use, then prove it with a short test
Adobe's help page for recording audio in Audition describes the mechanics: you select an input, arm the track you intend to record to, and capture a take, and a multitrack recording produces separate audio files tied to that session. That description comes from a single reading of the page and has not been re-checked since; it is documentation for one application, useful here because the same shape of steps exists in essentially every recorder.
What the documentation cannot tell you is whether the sound arriving at that input is the sound you think is arriving. That is worth thirty seconds of suspicion. Speak a line toward the microphone you meant to use and watch the meter. Then cover that microphone with your hand and speak again. If the meter still moves, something else is being recorded — a laptop's built-in microphone, a second channel on the interface, a phone left open on the desk. This happens constantly, and it is much cheaper to find now than after you have recorded a dozen takes and labeled them all.
A moving meter only tells you that something is arriving. It says nothing about whether the voice is usable, whether the room is loud, or whether the person reading can hear themselves. So listen back before you go further. Put on headphones and play the test. What you are listening for is simple: your voice, no hum, no obvious room bloom, and no delay.
That last one deserves a name. If you are hearing yourself through the computer's monitoring path, the sound reaches your ears a little after you speak it. Noticeable delay makes almost everyone stretch or push to match the late version of themselves, and then you are judging a performance that was shaped by the monitoring, not by the copy. Direct monitoring from the interface usually avoids the round trip. If you cannot get it, take one earcup off and record without hearing yourself, then judge from playback. Do not conclude the copy is too long because a lagging monitor made the reader crawl.
Record phrasing, not a clock
Before you roll, give the sentence its situation in one or two lines. Not "put more energy in it" — that instruction produces volume, not meaning. Say who is speaking, to whom, and what the picture is doing: you are telling someone standing in your shop, holding a wheel, and the shot cuts to the wheel as you finish the thought. Readers who know the situation generally find the emphasis on their own, and it is usually better than the emphasis you would have assigned.
Then take the unforced pass first. This is the discipline in the whole exercise. If you already know the window is seven and a half seconds, you will produce seven and a half seconds, and you will have learned nothing about whether the copy belongs in that window. Get a read that sounds like a person saying something, then discover how long it turned out to be.
Leave air: a beat before the first word, a beat after the last, and real space between lines. You can shorten a gap between two lines in the edit. You cannot un-swallow a syllable, and you cannot add back the consonant you clipped because the take started on the first frame. The gaps are also the only give you have if the read turns out tight against the picture.
When the copy does not fit, record a marked alternative. Marked means the label says which kind of change it is. A different pace over the same words is a delivery alternative. Different words are a copy change, and they need a version number. Say so out loud on the take and write it in the file name, because in three weeks "the version where she says it faster" will not tell anyone whether the script changed.
The alternative you should not record is the one where you quietly compress the read to hit the window. That take will fit, and it will be useless, because the phrasing you were trying to test has been squeezed out of it.
A worked example: twenty words, two patterns, one trim
Everything in this example is invented to make the arithmetic checkable. No microphone was opened for this article, no performer was involved, and none of the durations below were measured. They are constructed.
The spot is a fictional twelve-second cut for Harbor Street Repair, a shop that fixes bikes instead of selling new ones.
The picture:
- 0:00.0–0:02.5 — a rusted chain on a bike leaning against a fence.
- 0:02.5–0:06.5 — hands truing a wheel.
- 0:06.5–0:09.5 — the rider rides off.
- 0:09.5–0:12.0 — logo card, music only. No voice intended here.
So the read has to be finished by 0:09.5. The copy is twenty words:
L1 Most bikes get replaced long before they're worn out. L2 We'd rather fix the one you already own. L3 Harbor Street Repair.
Pattern A, quick and even. Lead-in silence of 0.4 seconds, then L1 at 0:00.4–0:03.4, a short gap, L2 at 0:03.6–0:06.3, a short gap, L3 at 0:06.5–0:07.5. It ends at 0:07.5, clearing the window by two seconds. The landings are good — "worn out" arrives just as the hands come up on the wheel, and the shop name sits inside the ride-off. But two seconds of picture sit empty before the logo, and the read sounds like a person hurrying through a list.
Pattern B, unhurried, with a beat before "worn out." Lead-in of 0.4 seconds, L1 at 0:00.4–0:04.4, gap, L2 at 0:05.1–0:08.5, gap, L3 at 0:09.4–0:10.7. It ends at 0:10.7 — 1.2 seconds past the window. And the failure is specific rather than general: "Harbor" begins a tenth of a second before the logo card, and "Repair" lands entirely underneath it, over a card that was built to carry music and nothing else.
That is what you want a scratch track to produce: not "it's too long," but a single sentence naming which syllable is over the wrong image.
Now the copy change, recorded as a separate labeled take, still in pattern B. Trim L1 to "Most bikes are replaced too early" and L2 to "We fix the one you own." Fifteen words instead of twenty. Same lead-in, same gaps, L1 at 0:00.4–0:02.8, L2 at 0:03.5–0:05.9, L3 at 0:06.6–0:07.9. It ends at 0:07.9, inside the window by 1.6 seconds, with the shop name cleanly over the ride-off.
The cost is real and worth naming. "Long before they're worn out" was doing specific work — it named the waste. "Too early" is a general claim, and general claims are easier to ignore. If the worn-out image is the one that matters, trim only L2 instead: "We'd rather fix yours," sixteen words total, ending at 0:08.8. That fits, with under a second of margin. Under a second is not much, once a real performer's natural variation is in the room, and a take that clears the window by seven-tenths of a second is not a plan — it is a take that happened to fit that day.
One capture problem would change the setup partway through, so it belongs in the story. In the constructed session, the first test line — "Most bikes get replaced—" — would come back with a puff of air on the b of "bikes" and again on the p of "replaced." A puff like that is called a plosive, and it is a mechanical problem, not a performance one. The usual cause is a reader standing roughly ten centimetres from the microphone, straight in front of it, speaking across the grille.
The correction: raise the microphone so it sits slightly above the mouth and angles down toward it, and move back to about a hand-span, roughly twenty centimetres. Leave the gain where it was, and re-record the test line.
The change matters for the comparison. Patterns A and B both assume the same position and the same distance. If A had been taken at ten centimetres and B at twenty, you would be comparing two setups and calling it a comparison of deliveries. Whatever you change, change it before both takes and then leave it alone.
Labels, slates, and raw takes
A take you cannot identify is a take you will re-record. The naming scheme only needs to carry four things: the spot, the copy version, the delivery pattern, and one line about why the take exists. Something like:
HARBOR-12 / pic_v3_12s / v1 / A / even — fits with 2s spare, sounds impatient
HARBOR-12 / pic_v3_12s / v1 / B / unhurried — L3 "Repair" under logo card
HARBOR-12 / pic_v3_12s / v2 / B / trimmed L1+L2 — ends 0:07.9, clean landing
Note what the naming does: the version number carries the text change and the pattern letter carries the delivery change, so those two arguments never collapse into each other. Note also the picture version. A timing judgement is only valid for the cut you judged it against, and when the edit moves, every duration note you wrote expires with it.
A slate is the spoken label at the head of a take. If you use one, say it, then hold a few seconds of silence before the first word, so the label can never be mistaken for part of the timed passage. The alternative — a marker in the session and a clean audio file — is tidier if the take is going straight onto the picture without a trim.
Then leave the raw files alone. Do not punch a new performance over a take you decided to keep; that overwrites the evidence for your earlier decision, and the earlier decision is exactly what you will want to re-examine when the client asks why the shorter copy exists.
If someone else reads the scratch for you, agree what the recording is for before you roll. A file that stays in the session folder and a file that plays in a client presentation are different arrangements, and the second one is easy to slide into by accident.
Repair the setup, not the sketch
The defects worth fixing are the ones you can hear without being told to listen for them: clipping that crackles, puffs on the hard consonants, a room that makes the voice sound boxed in, and level jumps that mean the reader drifted toward and away from the microphone between lines.
Consistent distance matters more than perfect distance. A reader who stays roughly twenty centimetres away for every line gives you takes you can compare; a reader who moves gives you a level change that reads, wrongly, as a performance change.
When something is wrong, re-record the short passage. Processing a scratch read to rescue it burns the time you were trying to save, and it cannot restore intelligibility that was never captured — a clipped take does not become an unclipped one. The one piece of useful feedback processing gives you is negative: whatever you do to a scratch read tells you nothing about how the final voice will be treated, so do not build habits around it.
Fix the room and the microphone before you fix the performance. If a take is unlistenable, you cannot judge the phrasing inside it, and you will end up re-recording under time pressure with the decision already half-made.
What the track still leaves open
Hand over the takes with their questions attached. The useful scratch track is the one that lets someone else in the room say "B is better, but the copy is a second long" without having to guess which file they are hearing or what changed between them. Labels, separate versions, and a picture reference turn a pile of recordings into a decision.
Everything in it stays provisional. The performance is one person's read, the timing is one afternoon's timing, and the quality is a room and a hand-span of distance. None of that constrains the performer who eventually gets cast, and none of it settles the final spot's length.
Which is why the rushed take deserves the suspicion it gets. A read that fits because the reader sprinted is not a solution — it is a take that fit once, on a day when nobody had to sustain it twice. When the waveform ends at the right frame and the read sounds like a hostage statement, believe the read, and record a marked alternative instead.
Frequently asked questions
What is a scratch voice track for, and what does it not settle?
It exists to make one argument possible: whether the copy fits the picture. You record it so you, a director, or a client can hear where a line lands against a cut, and so that when the answer is not quite, a file shows which words changed. It is not an audition, not a placeholder to be defended, and not the first pass at the finished voice. It does not establish what a professional performer must deliver, and its duration does not fix how long the finished spot will be.
Which timing question should the take answer?
Three different questions get called does the voice fit: total length, landing, and comprehension. Total length can be settled with a stopwatch and a text-to-speech read. Record yourself when landing or comprehension is live, because those turn on where emphasis lands. Separate fixed copy, such as a product name, legal phrase, or already-approved line, from provisional copy, and write the split down.
How can I prove I am recording the intended input?
Speak a line toward the intended microphone and watch the meter. Then cover that microphone with your hand and speak again. If the meter still moves, something else is being recorded, such as a laptop's built-in microphone, a second channel, or a phone left open. A moving meter only says something is arriving, so listen back on headphones for your voice, no hum, no obvious room bloom, and no delay. Computer monitoring delay can make a reader stretch or push to match the late version of themselves; direct monitoring from the interface usually avoids that round trip.
Why record an unforced pass first, and how should alternatives be marked?
If you already know the window is seven and a half seconds, you will produce seven and a half seconds and learn nothing about whether the copy belongs in that window. Give the sentence a situation, get a read that sounds like a person saying something, and discover how long it turned out to be. Leave air: a beat before the first word, a beat after the last, and real space between lines. When the copy does not fit, record a marked alternative. A different pace over the same words is a delivery alternative; different words are a copy change and need a version number. Say the change on the take and write it in the file name. Do not quietly compress the read to hit the window.
What does the Harbor Street example show, and what capture defect changes the setup?
It is invented to make the arithmetic checkable. The picture ends its voice window at 0:09.5. Pattern A, quick and even, ends at 0:07.5 with two seconds spare but sounds hurried. Pattern B, unhurried, ends at 0:10.7, 1.2 seconds past the window; Harbor begins a tenth before the logo card and Repair lands under it. A trimmed 15-word copy ends at 0:07.9 inside the window by 1.6 seconds but loses the specific work of long before they are worn out. Trimming only L2 gives sixteen words and ends at 0:08.8, under a second of margin. In the constructed session, a plosive on the b of bikes and p of replaced would call for raising the microphone above the mouth, angling it down, moving back to about twenty centimetres, leaving the gain, and re-recording the test line. Change the setup before both takes and leave it alone.