Set a Pitch Sample's Audio Level: Peak Normalization Is Not Loudness Matching
Set a Pitch Sample's Audio Level: Peak Normalization Is Not Loudness Matching
You set three takes to the same peak. You checked the meters. Then you play them back to back and one of them is still almost inaudible next to the others, or one of them is tiring to listen to in a way the meters never warned you about.
Nothing went wrong. Peak normalization did exactly what it says. The problem is that "the same peak" is not the same thing as "the same loudness," and the gap between those two ideas is the whole subject of this article.
The short answer: a peak is the loudest single instant in a file. Loudness is what a listener's ear accumulates over the whole passage. Setting peaks equal sets a ceiling. It tells you nothing about how high the furniture inside the room is.
Two measurements, two different questions
A sample peak is the largest absolute value any sample reaches. It's a single number describing a single moment. A door click, a breath into the mic, a plosive on a P, a chair scrape, one bark of laughter — any one of those can be the peak of an entire take, and none of them is what your client is listening to.
An integrated loudness figure describes energy across the length of the selection, weighted to resemble how hearing actually works. It is a measurement over time rather than an instant, and the standard measurements behind it discount near-silence rather than letting quiet stretches drag the number around the way a plain average would. It's usually reported in LUFS, where 0 is full scale.
The useful number sitting between them is the gap: how far the loudness of a passage sits below its own peak. Sometimes this is called the crest factor or the peak-to-loudness ratio. A tightly controlled, even voice might have a gap of ten to fifteen units. A hushed performance with one loud laugh in it might sit twenty-five or thirty below its own spike.
Here is the part that resolves your problem. That gap is a property of the material. It is not a property of the gain you apply. Any constant gain you add moves the peak and the loudness by exactly the same amount, which means the gap doesn't budge. Peak normalization is a constant gain operation. So it can never close a gap. It can only shift the whole file up or down along a line where the gap stays fixed.
That's why your samples still feel uneven.
What peak normalization controls
Audacity's manual describes Normalize as applying a constant amount of gain to the selection, based on the selected maximum, so the highest peak lands on whatever target you set. Constant is the operative word. Everything in the passage moves together. The relative pattern of loud and quiet inside the take is untouched.
This has a consequence worth sitting with. If the loudest moment in a take sits twelve units above the body of the performance, then setting that take's peak to your target automatically places the performance twelve units below the target. You didn't choose that. The material chose it for you. And two takes with the same peak but different gaps will end up with their bodies at different levels — which is precisely the unevenness you heard.
The effect also carries options that aren't leveling. One checkbox removes DC offset, which shifts the waveform's center; that's a separate repair, not a gain decision. Another lets the channels be normalized independently, which gives each channel its own gain and therefore changes the relationship between left and right. If you leave a checkbox like that on and then describe the run as "peak normalized," your notes are incomplete. For an ordinary stereo pair whose balance you want to preserve — a music bed, a stereo room pair, a two-mic capture of one performance — keep the channels linked unless you have an actual reason not to.
Documentation for the two effects:
https://www.audacityteam.org/manual/effects/volume-and-compression/normalize/
https://www.audacityteam.org/manual/effects/volume-and-compression/loudness-normalisation/
Dialog wording and the values pre-filled in the boxes vary by version, so read the ones in front of you.
What loudness matching controls
Loudness Normalisation works from a measurement of loudness rather than a measurement of peak. It offers a perceived-loudness mode and an RMS mode, and both describe energy across the selection instead of the single highest sample.
The operation still applies one gain to the whole selection, just like peak normalization. What changed is how the amount of that gain was decided. So loudness matching does not flatten the loud and quiet parts inside your take. It does not turn a dynamic performance into a compressed one. It moves the whole take so its measured loudness lands on the target you declared.
It has channel options of its own, and the same warning applies: adjusting channels independently changes their balance.
Two things loudness matching does not promise.
First, it does not promise identical perception for every listener. Two files that measure the same loudness can still feel different in a room, on a phone speaker, through earbuds, in a car. A short spoken passage and a long music bed with the same reading will not land the same way. Measurement gets you into the right neighborhood. Listening decides whether you're at the right address.
Second, it does not tell you what the target should be. The number that happens to be sitting in the dialog box when it opens is a default. A default is not your recipient's requirement. If the recipient named a specification, use theirs. If this is a private pitch review with no stated specification, you are choosing a comparison target for your own purposes, and you should write down that you chose it and why. There is no standard for a private pitch, and adopting a broadcast or streaming figure as if it were one is an invented requirement.
What neither operation does
Neither one repairs a performance whose level wanders within a single phrase. That's a different job with a different set of choices, and normalizing repeatedly will not get you there.
Neither one fixes what's in the channels. If a voice is only in the left channel, both operations will faithfully preserve that mistake at a healthier level.
Neither one removes distortion. If the recording clipped when it was captured, normalization scales the damage; it doesn't undo it.
And neither one tells you what target is correct. That's a decision, and the rest of this article is about making it deliberately.
A worked comparison
The numbers below are invented round figures chosen to make the arithmetic visible. No audio file was produced, measured, or played for this article, and the values are not readings from anything. Run the same procedure on your own samples and expect different numbers with the same shape.
Suppose two takes. Take A is a hushed read with one loud laugh a few seconds in. Take B is a steady read at an even level with no spikes.
| Highest peak | Integrated loudness | Gap | |
|---|---|---|---|
| Take A | -3.0 dBFS | -27 LUFS | 24 |
| Take B | -3.0 dBFS | -18 LUFS | 15 |
That difference of nine units is what you're hearing. Same peaks, very different listening level.
Now peak normalize both to -1.0 dBFS. That's +2.0 dB of gain on each.
| Peak | Loudness | |
|---|---|---|
| Take A | -1.0 dBFS | -25 LUFS |
| Take B | -1.0 dBFS | -16 LUFS |
The peaks match. The gap is still nine units. The thing the listener noticed did not move at all. This is the failure mode from the opening of the article, reproduced in arithmetic.
Now try loudness matching — starting again from the two untouched originals, not from the peak-normalized versions above. Declare a target — say -18 LUFS, as an internal comparison target for this review, not as a standard.
Take B is already at -18 LUFS. No gain needed, peak stays at -3.0 dBFS.
Take A needs +9.0 dB to reach -18 LUFS. Its peak sits at -3.0 dBFS. Adding nine decibels of gain puts that transient at +6.0 dBFS: six decibels past the top of the scale. It clips.
So the declared target is not reachable on Take A without either clipping the transient or reducing it — and reducing a single transient is dynamic-range work, a different operation with its own deliberate set of choices, not a side effect you should let a leveling tool perform.
You could lower the target instead. Try -25 LUFS:
- Take A gains +2.0 dB, landing at -1.0 dBFS on the spike and -25 LUFS overall. It meets target, but it now sits a single decibel below full scale. Almost no room left.
- Take B loses 7.0 dB, landing at -10.0 dBFS and -25 LUFS.
That works arithmetically, and you should notice what it cost. Take B was pulled down seven decibels to accommodate a problem that belongs to Take A, and Take A is still leaning on a spike with one decibel of headroom.
The honest options are these. Match the loudness target and accept that the transient needs work first — a separate decision. Match loudness to a target Take A can actually reach with real room to spare, and accept that Take B comes down with it. Or don't match loudness at all: declare that this review is comparing content, note that the takes differ in level, and say so in the review notes rather than hiding it. Any of the three can be right. What makes it right is that you chose it on purpose and wrote it down.
Headroom and channel relationships
Matching a loudness target does not create safety. Neither does peak normalization — your target peak is a choice, not a guarantee.
Why leave room at all? Because the file will be played back through routes you don't control. A browser, a phone, a video editor with its own gain stage, a codec that resamples and can nudge a waveform slightly past where it started. A file sitting at 0.0 dBFS will clip on some of those routes. A decibel or so of headroom is cheap insurance, and how much you leave depends on where the file is going.
Then measure again after the gain change. Don't assume the result from the setting you typed in.
And keep the channel relationship deliberate. Two channels of one stereo performance are not two independent sources. If your material is a stereo music bed or a stereo pair on a room, the balance between those channels is part of the recording, and independently normalizing them rewrites it. If your material is genuinely two separate sources that happen to share a file, then per-channel treatment may be exactly right — but that's a decision you make, not a default you inherit.
Measure, listen, reopen the file
The procedure that holds up looks like this.
Decide what the file is for and who's listening, and write the target down before you open any effect. If the number came from the dialog box, say that. If it came from your recipient, say that. If you picked it for a comparison you're running yourself, say that too.
Look at the source before you change it. Find the peak. Find the loudness. Read the gap. Note where the loud moment sits and what surrounds it. Then ask the question that routes the rest of the work: is the unevenness between files or inside one? Between files is this article's subject. Inside one is a different job.
Choose the operation from the reason, not from habit. Peak normalization when you need a ceiling you can state. Loudness matching when you need several takes to be comparable in listening level.
Record every checkbox. DC offset removal, channel linking, measurement mode. Whatever you turned on or off, it goes in the note.
Apply the gain, then look at the peaks again.
Then listen to original, peak version, and loudness version in the same session, on the same playback route, at the same playback volume. Change nothing else between them. You're testing a specific set of questions: does anything distort? Does the channel balance survive? Does one take now feel like it's pressing forward while another recedes? And most importantly, does the information the pitch actually depends on — the line you need the client to hear — come through on all three?
Finally, reopen the exported file. Not the project. The file you're going to send. Check its level there, and check that stereo is still stereo and mono isn't. Exports carry settings people forget they set.
Keep the original untouched. Gain is reversible in principle right up until the moment you overwrite the source.
End where the decision actually is
Leveling a review copy is not about making your samples identical. It's about making the comparison fair, so that when the client reacts to something, they're reacting to your idea and not to your gain staging.
So end with three things stated plainly: the method you used, the target you used and where it came from, and confirmation that you inspected the file you're actually sending. If the operation you chose cannot satisfy the sample's own constraints — if reaching the loudness target means clipping the transient, or if the material varies so much inside a single phrase that any one gain is a compromise — then the finished result is a stated limitation and a separate dynamics decision. Not a number you forced onto the file because the dialog box offered it.
Frequently asked questions
Why can three takes at the same peak still sound uneven?
A peak is the loudest single instant in a file; loudness is accumulated energy over the passage. Peak normalization applies constant gain so the highest peak lands on a target, which sets a ceiling but says nothing about how high the body of the performance sits. The gap between loudness and peak—sometimes called crest factor or peak-to-loudness ratio—is a property of the material. Any constant gain moves peak and loudness by the same amount, so the gap does not budge. Takes with the same peak but different gaps end up with their bodies at different levels.
What does peak normalization actually control?
It applies a constant amount of gain to the selection, based on the selected maximum, so the highest peak lands on whatever target is set. Everything in the passage moves together, and the relative pattern of loud and quiet inside the take is untouched. If the loudest moment sits, say, twelve units above the body, setting that take's peak to a target places the body twelve units below the target. The effect may also offer DC offset removal, which is a separate repair, and independent channel normalization, which changes the left-right relationship.
Does loudness normalization flatten dynamics or tell me the right target?
No. Loudness Normalisation measures energy across the selection—perceived loudness or RMS—but still applies one gain to the whole selection. It does not turn a dynamic performance into a compressed one. It also does not promise identical perception for every listener or playback route, and it does not tell you what target to use. The number in the dialog is a default, not the recipient's requirement. If a recipient named a specification, use theirs; if this is a private pitch review, you are choosing a comparison target and should write down that you chose it and why.
What if a loud transient prevents reaching the loudness target?
In the invented comparison, Take A needs +9 dB to reach -18 LUFS, but its peak at -3.0 dBFS would rise to +6.0 dBFS and clip. Reducing a single transient is dynamic-range work—a different operation with its own deliberate choices—not a side effect to let a leveling tool perform. The honest options are: match the target and accept that the transient needs work first; choose a target Take A can actually reach with real headroom and accept that Take B comes down with it; or do not match loudness at all, declare that the review is comparing content, note that the takes differ in level, and say so in the notes. Any can be right if chosen deliberately and written down.
What leveling procedure holds up for a pitch sample?
Decide what the file is for and who is listening, and write the target down before opening any effect. Look at the source: peak, loudness, gap, where the loud moment sits. Ask whether the unevenness is between files or inside one; between files is this article's subject, inside one is a different job. Choose the operation from the reason—peak normalization for a stated ceiling, loudness matching for comparable listening level. Record every checkbox, including DC offset removal, channel linking, and measurement mode. Apply gain, then look at peaks again. Listen to original, peak version, and loudness version in the same session, on the same route, at the same volume. Reopen the exported file and check its level, channel layout, and that stereo is still stereo and mono is not. Keep the original untouched.