Skip to content

An Uneven Voice Track: Ride Clip Gain, Compress It—or Record It Again?

Advertising

An Uneven Voice Track: Ride Clip Gain, Compress It—or Record It Again?

The reflex when a voice read comes back uneven is to put a compressor on it. The reflex isn't stupid. Compression is the tool that exists for exactly this shape of problem. It's just usually reached for before anyone has asked the question that decides the answer.

The answer, briefly. Whole-track gain — peak normalization, loudness normalization, a single gain move across the file — applies one change to everything, and one change to everything cannot repair a difference inside a take. The internal repair has three routes, and the diagnosis picks between them. One moment has moved and nothing else about it changed: ride it by hand. The level wanders across the read in small steps nobody meant: compress it. The microphone no longer carries what was said — clipped, blunted, buried: record it again. Three routes, three different costs.

A read with two different problems

The example below is invented. No audio was recorded or processed to write this, and none of the settings mentioned were tried on a real track. It exists to make the decision structure visible. You run the diagnosis on your own file.

Two sentences, one pass, one microphone, one performer. The first: "It's the same every morning." The second: "We built the thing that ends it." In the take, the first sentence lands quiet, because the performer let their level fall for it — same distance to the mic, no turning away. "Morning" is soft, but the consonants are all still in it. The second lands loud, because the performer leaned in and drove "ends it" — and that push is the reason the director approved the read. The loud moment is intact; it is not clipped.

Two sentences, two different reasons for the gap between them. The director wants the second one preserved. Everything that follows is about not confusing them.

What actually moved

Before choosing a tool, listen for cause rather than amount. Four checks, roughly in this order.

Look at the push. If the tops of the loud waveform are flat — clipped — then the detail is gone and neither gain nor compression will return it. That changes the whole decision, so establish it first. Assume, as in the example, that the push is loud but intact.

Then test the quiet part for more than level. Bring that phrase up provisionally until its loudness matches the sentence next to it, and listen to the consonants. If, once matched, it still sounds like the same person talking from further away — the sibilants and the t sounds blunter, more room around the voice, less edge on the vowels — then the microphone captured a different sound, not merely a lower level. Gain makes loud what was captured. It doesn't restore what wasn't. This is the single most useful check in the whole procedure, because it divides "wrong level" from "wrong recording," and those two have different remedies. In the example the check passes: lifted, the first sentence is the same voice at the same distance, only louder.

Locate the boundaries. Does the level change inside a word, between words, on a breath, in a pause, between sentences? Where the step falls decides how much a repair will cost in audible artifacts.

Check the room under the quiet part. If the noise floor is audible there, every upward move takes it up too.

That gives you the taxonomy the rest of the article works from. The variation in front of you is one of three things: useful performance, meant and to be protected; a level change that moved nothing else, the same voice in the same room, only quieter or louder on the meter; or a capture change or damage — position, distance, off-axis, a clipped peak, a noise floor you can't afford to lift — meaning something the microphone no longer carries and no amount of gain can put back. The categories aren't exclusive, and one take can hold all three, but the route follows the dominant one.

And note the asymmetry in the example: only one of the two sentences is a defect. The push is not a problem to be evened out. Any setting chosen to "make the read more consistent" is a setting aimed at the emphasis.

The repair that follows meaning

Manual clip gain and automation are the same idea at different resolutions: you decide, region by region, that a phrase should sit higher or lower than the raw take put it. That sounds like the clumsy option. In this situation it's usually the most conservative one, because it responds to meaning. An editor can hear that a loud phrase was intentional and leave it alone. A compressor cannot hear anything.

Three things make the difference between a ride that works and one that fights you.

Match to the body of the read, not to the peak. Bring the quiet sentence up to the level of ordinary speech around it. If you bring it up to the push's level, you haven't repaired an accident — you've deleted an emphasis, using a tool that at least would have taken the blame.

Ride only what is wrong. In the example that means the first sentence and nothing else. The loud half of the problem isn't half the problem.

Place the boundaries in pauses or on breaths, with a short fade rather than a hard corner. A step in the middle of a sustained vowel is audible as a fault; a step across a breath is usually inaudible.

Then listen for what you paid. Raising a phrase raises the room under it by exactly the same amount, so the noise floor now steps up during that sentence and back down after it. On headphones, in a quiet pitch video, that step is an artifact even when the voice itself sounds fine. Whether it matters depends on the bed and the playback level, which is why the comparison later happens in context, not in solo.

What riding cannot do: restore dullness, recover clipped detail, or rescue a phrase whose words are unintelligible at any level. The first sentence in the example needs none of those. If a quiet phrase needed any of them, gain was never going to be the answer.

The process that follows level

Compression is easier to describe than to decide with. Above a threshold, gain is reduced by an amount set by the ratio; attack and release govern how quickly it responds and recovers; makeup gain restores overall level. The Audacity manual documents those controls on its compressor page, which carried an update date of September 14, 2026 in the version of its HTML text this piece was written against. Control definitions, though, don't tell you what a setting will do to a performance.

Here is where it points in this example. Suppose you set the threshold eight decibels below the push — low enough to catch it — at a 3:1 ratio. The control's arithmetic brings that push down by roughly 5.3 dB: the excess above threshold, multiplied by one minus the reciprocal of the ratio. The quiet first sentence sits below threshold, so on its own the compressor never touches it.

That is the shape of the problem. To lift the first sentence you need makeup gain, and makeup gain raises everything below threshold by the same amount — the quiet speech, the breaths, the room tone under all of it — while the push ends up approximately back where it started. The two sentences are now about 5.3 dB closer together than the director wanted, and the noise floor under the first is 5.3 dB louder. That's the arithmetic of the control and not a measured result from any file; attack, release and the material will move the numbers. The direction is what matters, and the direction doesn't change.

Which leaves the threshold as the real decision, and every sensible placement fails here for a different reason. Set it above the push and nothing happens at all — the compressor is a no-op for this problem. Set it between the push and the quiet speech and the process works against the emphasis while doing nothing for the accident. Set it below the quiet speech and everything is compressed: the read flattens as a whole, and makeup gain brings the room up with it. That third setting is where the complaint about compression sounding "even in the wrong way" comes from.

The timing controls reach into things you care about too. A release long enough to still be holding gain reduction after a loud passage will pull down the quiet phrase that follows before it recovers — a level change in the opposite direction from the one you wanted. An attack fast enough to catch the first milliseconds of a consonant reduces gain during that transient, which can take the edge off the word's opening. When you audition the result, listen past "is it steadier now" to phrase openings, phrase endings, breaths and pauses. Those are the places the processing shows up.

So compression is the right tool when the variation is diffuse: a performer drifting across a paragraph, dozens of small bumps, no single move that was meant. It's the wrong tool when the variation is one deliberate gesture and one accident, because the mechanism has no way to prefer one over the other. Reverse the example — imagine the second half had drifted loud without anyone intending it, and the quiet speech were the body of the read — and compression becomes the first thing worth trying. Same tool, opposite verdict, because the cause changed.

When to record it again

Rerecording is the only route that changes what the microphone captured, which makes it the only route that cures clipping, off-axis dullness you can't accept, a distance change large enough to alter the balance of direct voice to room, or a noise floor you can't afford to raise.

The cost is the performance. A new take is not the approved take with a technical repair applied; it's a second performance, and it will differ in ways that aren't about level. The push may come back with a different weight. The pace may shift a step. That risk is real and it should be named out loud before anyone books the room.

There's also a choice about scope. Patching one sentence costs the performer less, but it puts two performances inside one read: a timbre and pace match at the join, a level match that may or may not hold, and a strong temptation to paper over the seam later with more gain. Rerecording the whole read gives you one consistent performance, and it may simply relocate the unevenness somewhere new. Decide by asking what was actually approved — one sentence, or the read as a whole.

And sometimes the honest conclusion is to keep the flaw. If the quiet sentence is intelligible, sounds like the same person in the same room, and reads as an aside rather than a mistake, a small level nudge may be the entire repair. That's a judgment to make with the director, not in the waveform.

How to tell which one won

Whatever you build, duplicate before you touch. The raw take stays untouched, and each route lives on its own copy, because the original is the only thing you can return to, and it's also what you'll need if the performer becomes available again.

Then level-match before you judge. Louder sounds better, reliably and unhelpfully. Peak normalization sets a single gain based on the loudest sample in the selection; loudness normalization measures perceived loudness instead. For an A/B between a compressed and an untouched version, the loudness-based comparison is the closer one, because compression with makeup changes the relationship between peaks and average level — two versions can share a peak and sit at noticeably different perceived loudness. Neither operation fixes internal unevenness; that's a separate job. The normalize dialog's optional correction for a waveform that doesn't sit on the center line is a third job again, and not a level repair at all.

Then compare in three passes:

  1. Solo, one dimension at a time. Consonants at phrase openings. The shape of the push. The room tone in the pauses. The breaths.
  2. Solo, whole take. Does this still sound like a person making a point, or a file being managed?
  3. In context. The actual sketch, at the volume the pitch will be heard at. Solo flatters repairs that add noise; the bed and the real playback level are where a noise step gets caught, or doesn't. If music will cover most of the read, some of that noise may be masked — but verify it in the bed rather than assuming it.

Record which defect each version addresses and which it leaves behind. The purpose isn't paperwork; it's that a decision you can re-open is worth more than a decision you have to defend. And there is no preset to hand anyone here. A setting follows a diagnosis, and anyone offering you the setting first has skipped the diagnosis.

Noise reduction, incidentally, is a different operation with its own cost to the voice, and it isn't the tool for a level difference. It belongs to a different question.

The least destructive route

Three conditions, three answers, and each keeps the rejected versions visible.

If one phrase moved in level alone, the room that comes up when you lift it is acceptable in the mix, and the push was approved as performed — ride it. Preserve the emphasis by not touching it.

If the level drifts all over the read and no single move was meant, a light compression is a smaller intervention than a hundred hand edits, and the arithmetic that makes it wrong for this example makes it right for that one.

If the information is gone — clipped, blunted, buried, or noisy enough that lifting it would bring the room with it — record it again, and accept that you're trading a technical problem for a performance variable.

There's no fourth route that recovers what wasn't captured, and no processor that knows which loud moment was a decision. Evenness was never the goal. The goal is a track where every word is audible and the emphasis still sounds intended, and where those two pull against each other, the emphasis is the part you can't get back.

Frequently asked questions

How do I choose between riding clip gain, compressing, and rerecording an uneven voice take?

Diagnose the cause first. If one moment moved in level alone and nothing else changed—same voice, same distance, same room—ride it by hand and protect any intentional push. If the level wanders across the read in many small steps nobody meant, light compression is the smaller intervention. If the microphone no longer carries what was said—clipped, blunted, buried, or a noise floor that cannot be lifted—rerecord. The routes have different costs.

Why can't whole-track gain or normalization fix an uneven take?

Whole-track gain applies one change to everything. It cannot repair a difference inside a take. Peak normalization sets a single gain based on the loudest sample; loudness normalization measures perceived loudness. Both still apply one gain to the selection, so neither restores clipped detail nor fixes internal unevenness.

What checks should I run before choosing a repair?

Look at the loud waveform for clipping, because clipped detail is gone. Lift the quiet phrase provisionally to match neighboring speech and listen to consonants: if it still sounds like the same person at the same distance only louder, the problem is level; if sibilants and t sounds are blunter and the room is more present, the microphone captured a different sound. Locate where the level changes—inside a word, between words, on a breath, in a pause, between sentences. Check the noise floor under the quiet part.

Why can compression work against an approved emphasis?

In the example, the loud push is intentional and the quiet first sentence is the accident. A threshold between the two makes the compressor work against the emphasis while doing nothing for the quiet phrase. Makeup gain then raises the quiet speech, breaths, and room tone, while the push returns near its starting level—reducing the intended contrast and raising the noise floor. Timing controls can also pull down a following quiet phrase or soften a consonant attack. Compression suits diffuse variation, not one deliberate gesture plus one accident.

How should I compare versions before deciding?

Duplicate before touching anything and keep the raw take untouched. Level-match before judging, and for compressed versus untouched material use a loudness-based comparison, because compression with makeup changes the relationship between peaks and average level. Compare in three passes: solo on one dimension at a time; solo as a whole take; then in context at the volume the pitch will be heard. Record which defect each version addresses and which it leaves behind.

More in Advertising Browse all articles