Skip to content

Make Room for Voiceover Without Flattening the Music: Manual Fades or Generated Ducking?

Advertising

Make Room for Voiceover Without Flattening the Music: Manual Fades or Generated Ducking?

There is a reflex that costs more than it looks. A line of copy sinks under a music cue, someone grabs the music fader and pulls the whole cue down four or five decibels, and the words come through. The cue is now a smooth grey surface with no shape. The thing the music was doing—the entrance, the lift, the turn—has been paid for out of the one resource it needed.

The opposite reflex is faster. Click the music, switch on ducking, generate, move on. Now the music breathes in and out on every intake of air, and the syllable the spot depends on gets covered by a return that arrives a beat early.

Both failures skip the same step. Before you decide how to lower the music, decide what has to be intelligible and what has to survive.

That decision is the work. The envelope is its paperwork.

What follows is a way to make that decision in a few passes: separate the two signals, mark what must be heard on each, build a deliberate manual envelope, generate a second envelope and refuse to treat it as finished, then compare the transitions rather than the average level. The walkthrough uses a small invented spot. Nothing in it has been recorded, generated or listened to here, and the tool behavior it relies on comes from a documentation page rather than a session. That distinction runs through the whole method, so it is worth keeping in view while you read.

Separate the competing signals and intentions

A ducking envelope is a relationship between two audio roles, not a property of a track. The voice is the driving signal; the music is the attenuated one. If both live inside a single mixed file, no amount of control adjustment will separate them, because the tool has nothing to drive and nothing to attenuate. Get them onto separate tracks with correct roles first, keep each independently adjustable, and keep an unducked version of the sketch somewhere you can play it. That reference is not sentimentality. It is the only honest way to say later what you changed.

Then make a distinction that automatic processing cannot make for you. Music masking speech and an uneven voice recording are different problems with different levers. If the voice's own dynamics wander—one phrase sitting well under its neighbor, a take that starts loud and drifts—lowering the music buys intelligibility by spending the music, and the quiet phrase is still quiet against its own surroundings. That belongs to clip gain, compression or another take. Keep it out of the ducking decision, or you will reach for a cross-track tool to solve an in-track problem and wonder why the result sounds expensive and still wrong.

Now mark two things.

On the voice track, find the phrase that must land. Not the loudest one, and not every phrase. Usually one, maybe two, and they are the ones carrying the claim.

On the music track, find the event that carries the idea: an entrance, a change of texture, a harmonic turn, a payoff. Figure out where it falls against the speech. Sometimes it sits comfortably in a gap. Sometimes it collides with the indispensable phrase, and then you have to choose which one owns that instant. Making that choice explicitly is the difference between a mix and an accident.

Here is the invented spot, with its particulars fixed so the rest of this article can refer to them. It is a twenty-two-second treatment sketch, fictional, and not built.

The voice is one take of spoken copy: "We number every batch by hand." A breath. "That's slower than a machine." A short interjection: "Still is." A pause. Then the closing phrase: "The number is the point."

The music is a simple original cue: a held low line, and an upper line that enters with the closing phrase. That entry is the event. It is what the spot is spending its music on.

Three versions come out of this: an unducked baseline, a manually shaped envelope, and a generated one, plus a corrected copy of the generated one.

Build a deliberate manual version

A manual envelope has three parts, and it is worth naming them because most rushed fades only do the first. There is the move down, the floor under the words, and the return. A reduction that just descends and stays descends forever.

Place the move around the actual phrase rather than the whole cue. That sounds obvious and is the most common thing to get wrong, because lowering the entire cue is one gesture and shaping one phrase is four. If the phrase to protect is eight seconds into the spot, the first seven seconds should sound the way the music was written.

The floor has no universal number. How far down you need to go depends on what the music is doing in that instant. A sparse pad and a busy mid-range figure can require very different amounts of room to let the same syllable through. Set the floor while listening to the word, not while looking at the meter.

The shape of the return is where manual work earns its keep. Keyframes let you go down quickly and come back slowly, or the reverse, and asymmetry is most of what makes a level change sound intended rather than triggered. Two useful archetypes turn up constantly around a short pause. The first holds the reduction through the pause: the voice's thought stays continuous, and you spend music to get that continuity. The second lets the music rise in the pause and settle again: the music stays alive, but the return has to land somewhere that is not the onset of the next word, or the first syllable arrives under a moving target.

The return is also where manual fades most often go wrong. If the music comes back up while the last word is still ringing, that word gets covered by the very gesture meant to release it. Return after the word, not after the waveform technically crosses a line. If the next phrase is far enough away, you can let the music come back early and enjoy it.

Run that against the invented spot and the manual build has a clear job. The breath between the first two sentences is not a place the music should react to; it is a place where the voice is not saying anything and the cue can breathe. The interjection "Still is." is short, semantically load-bearing, and awkward: two syllables is long enough that a full reduction makes it sound heavy, and short enough that any envelope with a lazy return will not come back up before it is over. The closing phrase is the collision. The upper line enters with it, so the envelope has to be out of the way for the phrase's onset and back near baseline while the entry is still sounding. That is a narrow window, and it is the transition the whole decision will turn on.

Generate a second version without treating it as finished

The documented route in Premiere, as the Adobe page on automatic ducking describes it, runs roughly like this: classify the music as music and the voice as dialogue in Essential Sound, select the music, enable Ducking, choose what it ducks against, adjust the ducking controls, and generate keyframes. The result is an editable gain envelope, and the page warns that regenerating the keyframes overwrites your manual changes.

That page, "Automatically duck audio," was updated January 7, 2026; it was inspected September 18, 2026. It is a documentation observation. Nobody in this article has generated a keyframe or heard a ducked cue. Read the control names off the page in your own version rather than trusting a list here, because these controls move between releases, and note the version you used alongside your project.

The controls divide into four questions, whatever they are labeled: what triggers the reduction, how deep it goes, how quickly it moves, and how long it takes to come back. The generator answers all four mechanically. What it cannot know is whether a sound is a word.

So listen for the events that are not speech. Breaths, lip noise, a short interjection, a "mm" of assent, room tone that got caught by a gate. On the invented spot, the breath between the first two sentences is the obvious candidate: it sits exactly where the music could rise, and an envelope that treats it as dialogue will write a dip there instead. The interjection is the second: two syllables may be enough to trigger a full reduction and not enough to let the return finish before the next phrase, which is how an aside ends up costing the space around it. A very fast return has its own failure, jumping the music up between syllables so a single word gets a stutter of music under it.

Then listen for the gaps. A pause makes the music come back and go down again, and if the speech is broken into short fragments, the result pumps. The unducked baseline is the check: if a level move happens where the voice is silent, it is movement for its own sake.

The generated envelope is editable, which is the useful part. Preserve a copy before you touch it, because that overwrite warning has a practical consequence: changing a ducking control and regenerating discards the keyframes you hand-adjusted. If you tune the controls after correcting, you are starting the corrections over. So decide early whether that pass is the one you are keeping, duplicate the sequence or track if it might be, and correct afterward. And do not treat the raw generated result as finished. It is a survey of where the speech is, with a beginning-of-term estimate attached to each event.

Compare the transitions, not just the average level

To compare fairly, hold everything else still: same voice take, same cue, same playback system, same monitoring level across all three versions. The monitoring level matters more than it seems. Played quietly, the small events that make generated ducking sound mechanical—the breath dips, the interjection holds—are the first thing to disappear, and you will approve a version you cannot hear. Played loudly, everything sounds like it is pumping. Pick a level you can work at, note it, and don't move it between versions.

Then listen at the boundaries rather than under the loudest word. The onset of each phrase: did the music get out of the way in time, or is the first consonant sharing space with a descending envelope? The end of each phrase: did the music come back too soon and cover the tail? The pause: did anything move where the voice wasn't speaking? And the musical event: what state is the music in at the moment it matters most? That last question is the one the average level of a whole cue will never answer.

What you are counting is the number of gain moves and whether each one earns its place. A version that protects meaning with fewer distracting changes is the one to keep. This is deliberately not a decibel target or a fade length. A recipe would be someone else's decision wearing your spot's clothes.

If neither version works, the envelope may not be the problem. A breath that cannot be afforded under the words, a sibilant fighting the top end of the cue, a level jump between two takes—those are voice-side repairs, a different job with different tools. Rerecording or repairing the voice is a legitimate outcome of this comparison, and a better one than spending the music to hide it.

One boundary worth drawing around the whole exercise: a treatment preview is not a finished broadcast mix. What you are choosing here is an editable relationship between two tracks. Loudness delivery, stem exports and the approval chain are separate operations with their own specifications, and none of them are settled by the shape of a ducking envelope.

The invented spot decides on one transition: the return under the closing phrase, where the upper line enters. The envelope has to be low enough for the phrase's onset to read and back near baseline while that entry is still ringing. A uniform automatic release, working from a depth and a duration that were set before anyone heard the phrase, is unlikely to produce that shape on its own. Correcting the generated pass to get it is possible, but the correction is the same move you would have made by hand, performed inside a file that can be regenerated out from under you. That is an argument for shaping this passage manually from the start, not a verdict on generated ducking in general; passages with many speech events scattered across a long cue are where a generated starting point does its best work.

Whatever you choose, write down the reason next to the transition that justified it, and note the software version you tested. Keep the unducked source and the corrected envelope together. Six weeks from now, when the music changes or the read is replaced, you will want to know what you altered and why, rather than guessing from a waveform that has already been printed.

None of that has been done here. The spot is a construction, the comparison is a procedure, and the ducking behavior comes from a page rather than a session. Build it, listen at those boundaries, and the choice will be yours rather than an inherited default.

Frequently asked questions

What should be decided before lowering music under voiceover?

What has to be intelligible and what has to survive. Pulling the whole cue down flattens the music's shape; automatic ducking can breathe on every intake and cover a syllable. The envelope is paperwork for that decision.

What are the three parts of a deliberate manual ducking envelope?

The move down, the floor under the words, and the return. Place the move around the actual phrase, set the floor by listening to the word rather than a meter, and shape the return asymmetrically—often returning after the word, not while it is still ringing.

What can generated ducking not know, and what is the correction risk?

It can't know whether a sound is a word; it may treat breaths, lip noise, short interjections, or room tone as dialogue. The generated envelope is editable, but regenerating keyframes overwrites manual changes, so preserve a copy and correct afterward rather than treating the raw result as finished.

How should manual and generated versions be compared?

Hold the same voice take, music cue, playback system, and monitoring level, then listen at boundaries: phrase onsets, phrase ends, pauses, and the musical event. Count distracting gain moves and whether each earns its place, rather than comparing average level.

What if neither the manual nor generated ducking works?

The problem may be voice-side—an unaffordable breath, a sibilant fighting the cue, or a level jump between takes—which is a different job with different tools. Also, a treatment preview isn't a finished broadcast mix; loudness, stems, and approval are separate.

More in Advertising Browse all articles