Skip to content

Pitch Lip-Synced Dialogue When Voice and Performance Are Made Separately

Advertising

Pitch Lip-Synced Dialogue When Voice and Performance Are Made Separately

When a voice is recorded in one room and the body that appears to speak is captured or drawn somewhere else, the treatment has one job beyond describing the shots: say who the audience should believe is speaking, and say how closely the voice and the body are meant to agree. "Her lips match the track" is not a description of a performance. It's a prediction about a technical result, written before anyone has heard the line or seen it land on a face. It tells a producer, an editor, and a performer almost nothing about what to protect.

The version that works is narrower and more interesting. Name the apparent speaker. Name the feeling of the exchange. Then pick out the two or three places where agreement genuinely matters, and be just as clear about the places where the voice and the body are supposed to pull apart.

Start with who seems to own the voice

Before any line lands, the audience is answering a question: whose voice is this? Most scenes with separately made speech fall into one of three answers.

The visible person owns it. Voice and body are presented as one continuous self, and the audience is not invited to think about the join. This is the case worth writing precisely.

Someone off screen owns it, and the audience knows it. A narrator, a letter read aloud, a memory, a voice with no body in frame. Nobody expects a mouth to move, and the whole question of synchronization changes shape. That's a different proposition, about a speaker's relationship to the film rather than a body on screen carrying the words.

A constructed speaker owns it. An animated figure, a puppet, a mouth on a billboard. Still a body relationship, just a drawn or built one, and it comes with its own conventions about how much articulation is enough.

There's a fourth answer that treatments sometimes arrive at by accident. The voice belongs to nobody the audience can identify, and they read the scene as either a mistake or a narrator they were never introduced to. Sometimes that's exactly the idea. Often it isn't, and it happens because the treatment specified the words and left the attribution to chance.

Which means the register of the effect belongs in the treatment too. Is this meant to feel ordinary, broad and theatrical, or knowingly artificial? Those aren't post-production moods. They change what the performer does on the day and what an animator draws, so they have to be decided before the work is commissioned rather than discovered in a mix.

A short scene, two ways

The scene below is invented. Nothing here has been recorded, animated, or watched. Treat it as a diagram.

A kitchen, late. Dev pulls a baking tray from the oven. The cookies on it are black discs. He sets the tray on the stovetop and looks at Marisol, who is leaning against the counter with a plate and a spatula beside her. She says: "That was the plan."

Version one, aligned. The recorded line has a small breath before "That," a catch on "was," and a pitch that settles downward on "plan." It is a line trying to sound settled and not quite managing. On screen, her mouth shapes the words as they arrive. The smile comes a half-beat after the last one. Her eyes go to Dev, then to the tray, then back. Nobody speaks for about two seconds. Dev's mouth starts to open and stops. Marisol reaches for the spatula and moves one cookie toward the plate as though she were serving.

Here the timing is exact: the strike of "That" and the closed-lip closure on the "p" of "plan" land on the visible mouth. The delivery reaches for composure and falls short of it, and the face does the same thing a half-beat later. Voice and bearing agree. Aligned, in this version, covers both: the sound lands on the visible mouth, and the voice and the body are saying the same thing. She isn't covering for herself. She is caught, and everyone in the room can see it.

Version two, divergent on purpose. Same words, same mouth shapes, same body. Only the delivery moves: bright, quick, level or slightly rising, no audible breath, pitched as though she expects him to agree with her. Now the words are confident and the person isn't. That gap is the joke. She is arguing with her own kitchen. The placement is still exact; what has come apart is voice and bearing.

For that reading to survive, two things have to hold. The divergence has to live in the delivery, not in the sync. If the mouth stops shaping the words, the audience stops hearing a person and starts hearing a track. And the face has to stay legible as uncertain rather than blank: the smile still late, the eyes still moving. If her body goes confident as well, with a quick smile and easy shoulders, the two versions collapse into one and there is no gap left to notice.

There is an outer edge. Push the divergence far enough — bright vocal, face entirely still, no reaction to Dev, eyes fixed on the tray — and the attribution changes. The audience stops hearing Marisol cover for herself and starts seeing a voice laid over her picture. At that point the scene's meaning shifts from a person lying about cookies to a problem with the film. If that shift is the idea, it's a different idea, and the treatment has to say so in those terms.

Four things that can agree, and don't have to agree together

Once you have a scene like that in hand, the distinctions get easier to hold.

Placement in time. Which sound lands on which visible mouth movement. This is what most people mean by lip sync, and it's the narrowest of the four.

Vocal identity. Who the audience thinks the voice belongs to. A different age, accent, or register changes this, and it can change it enough to alter the idea. If a scratch voice is later replaced by one that sounds like a different person, the joke in version two may not survive the swap, because the gap was never only about words and timing.

Bearing. What the rest of the body does: posture, gaze, hands, breath, the beat before the line and the beat after it. In both versions above, the late smile is bearing, not synchronization; in the second, the distance between it and the voice is doing most of the work.

The listener. Dev's non-reaction is an instruction to the audience about how to take the scene. If he laughs, he tells them it's a joke. If he doesn't, they have to decide for themselves whether Marisol landed the cover story.

A treatment can lock placement and leave bearing loose, or lock bearing and leave placement to the edit. What it can't do is leave all four to the phrase "we'll make it match," because that phrase doesn't tell a performer where to put the hesitation, an animator which shape to hold, or an editor what to preserve when the timing is tight.

How much drift a viewer forgives before the illusion breaks is a real question with a real answer. It is also an empirical one, and this piece isn't going to invent a number for it. That's what the test below is for.

Pauses and the non-speaking beat

The two seconds after the line are not dead air. They are where the audience decides what they just saw. If Marisol speaks again immediately, the cover story gets accepted without a moment of doubt and the scene flattens into ordinary banter. The silence is a scheduled element, and its length is a decision on the same footing as the words.

The same goes for what the body does while nobody is talking. The spatula, the plate, the cookie moved as though this were service — that's content. So is Dev's aborted sentence, which tells the audience to hold rather than to laugh. A treatment that lists only the spoken words and the mouth shapes has left out the performance and kept the transcript.

Say what gets made first

Separately made voice and image have a making order, and the order changes what the treatment needs to promise.

Voice first. A scratch or approved track exists before the camera rolls or before animation begins. The performer plays to the phrasing, so the breaths are available to copy and the timing can be built into the body. Two things belong in the treatment here: who records the track, and the fact that a scratch track is temporary. It's easy to slide from "the line is already recorded" to "we'll use the line we have," and that slide crosses a rights line nobody in the room may have agreed to.

Picture first. The visible performance exists, and the voice is recorded against it. Bearing stays unconstrained and real, because nobody was counting frames while it was performed. The cost is phrasing: the line may need to be shortened to fit the mouth that already moved, and breaths have to go wherever the picture leaves room. The treatment should say which words are fixed and which may be trimmed, because that's a decision someone downstream will otherwise make alone.

Voice first, with articulation deliberately pushed. Animation and some comic registers want mouth shapes larger than speech strictly requires. That's a divergence of its own, and if the treatment doesn't call it a style choice, somebody will helpfully correct it later.

Whatever the order, the treatment should distinguish fixed words from open ones, name the timing dependencies (the length of the pause, the moment of the look), and flag the decisions still outstanding. And it should not imply that a scratch performance, a likeness, or a recognizable voice becomes reusable because it has been recorded. Replacing a real person's voice, or reusing an existing one, is a permission and terms conversation that belongs well before the treatment circulates. That problem is not this one.

Plan the comparison as a viewing, not a map

The aligned and divergent versions have to be seen and heard together. A column of frame numbers, a row per syllable, a timing chart — these document an intention. They don't test it. Nobody has ever flinched at a spreadsheet.

When the comparison happens, the useful questions aren't only about sync. Does the mismatch read as intended, or as broken? That's a different question from whether the mouth matched. Ask a few people who said the line; if the answer changes between the two versions, the attribution has moved even though the picture didn't. Watch whether Dev's silence plays as a beat or as confusion about the film. And if the divergence lands as an error rather than a joke, revise the performance proposition instead of assuming a cleanup pass will bring the idea back.

What this piece can't tell you

The two variants above are written, not recorded. No performer has been asked, no animatic exists, and nothing has been watched or evaluated. The comparison described here is a plan for one test, not a report from one. Whether the mismatch reads as a joke is exactly the thing the test would establish, and it can't be established on the page.

The vocabulary is unsettled too. Lip sync, dubbing, ADR, voice replacement, localization — these overlap in ordinary use and carry different craft and rights implications depending on who's speaking and what shop they work in. Before a treatment puts a label on a real workflow, check with the post team which practice they actually mean.

And the neighboring problems stay neighboring. Once a performance is in the can, replacing its speech raises a preservation question about the take you already have, which is its own piece of work.

Write the paragraph that can be built

Here is what the aligned version looks like on the page, as an illustration rather than a record:

Dev lifts the tray out of the oven. Black cookies. He sets it down and looks at her. Two seconds. MARISOL, still leaning on the counter, says it as if the timing were hers: "That was the plan." Her mouth gives him the words. The smile comes a half-beat late. She holds his look, then reaches for the plate. He doesn't laugh.

Everything the people building this need is now visible. The voice is recorded separately and is hers. Timing is exact through "plan." The delivery and the bearing carry the same uncertainty, and the half-beat on the smile is what makes it legible. The pause is two seconds, not a suggestion. Dev's non-reaction is required, not optional coverage.

What must agree, and what may differ, is now a description of a relationship rather than a hope about a process. Whether the relationship holds is for the shoot, the animatic, or the cut to say. Matching is a decision. So is the mismatch. The treatment's job is to say which one you're making, and where.

Frequently asked questions

What should a treatment say when voice and visible performance are made separately?

It should say who the audience should believe is speaking and how closely the voice and body are meant to agree. 'Her lips match the track' is a prediction about a technical result, written before anyone has heard the line or seen it land on a face; it tells a producer, editor, and performer almost nothing about what to protect. Better to name the apparent speaker, name the feeling of the exchange, pick the two or three places where agreement genuinely matters, and be clear about where voice and body are supposed to pull apart.

What are the possible answers to whose voice the audience thinks it is?

The visible person owns it: voice and body are presented as one continuous self, and the audience is not invited to think about the join. Someone off screen owns it and the audience knows: a narrator, a letter read aloud, a memory, a voice with no body in frame; nobody expects a mouth to move. A constructed speaker owns it: an animated figure, a puppet, a mouth on a billboard, with its own conventions about articulation. A fourth answer can arrive by accident: the voice belongs to nobody the audience can identify, and they read the scene as a mistake or an unintroduced narrator. Sometimes that is the idea, but often it happens because the treatment specified the words and left attribution to chance.

In the invented kitchen scene, what distinguishes the aligned version from the divergent version?

The words, mouth shapes, and body can be the same. In the aligned version, the recorded line has a small breath before 'That,' a catch on 'was,' and a pitch that settles downward on 'plan'; the delivery tries to sound settled and does not quite manage it. Her mouth shapes the words as they arrive, the smile comes a half-beat after the last word, and voice and bearing agree: she is caught, and everyone can see it. In the divergent-on-purpose version, the delivery is bright, quick, level or slightly rising, with no audible breath, as though she expects him to agree. The words are confident and the person is not; that gap is the joke. For that reading to survive, the divergence must live in the delivery, not in the sync, and the face must stay legible as uncertain rather than blank. Push the divergence far enough and the attribution changes from a person lying about cookies to a voice laid over her picture.

What four things can agree, and why does it matter that they do not have to agree together?

Placement in time: which sound lands on which visible mouth movement, the narrowest sense of lip sync. Vocal identity: who the audience thinks the voice belongs to; a different age, accent, or register can change the idea, and a scratch voice replaced by one that sounds like a different person may not preserve a joke. Bearing: what the rest of the body does, including posture, gaze, hands, breath, and the beats before and after the line; the late smile is bearing, not synchronization. The listener: Dev's non-reaction instructs the audience how to take the scene; if he laughs, he tells them it is a joke, and if he does not, they decide for themselves whether the cover story landed. A treatment can lock some and leave others loose, but it cannot leave all four to 'we'll make it match,' because that tells a performer, animator, or editor nothing about what to preserve.

What changes depending on whether voice or picture is made first, and what does the article say about testing and limits?

Voice first: a scratch or approved track exists before camera or animation, the performer plays to the phrasing, breaths can be copied, and timing can be built into the body; the treatment should say who records the track and that a scratch track is temporary, because sliding from a recorded line to using it crosses a rights line. Picture first: the visible performance exists and voice is recorded against it, so bearing stays unconstrained and real, but phrasing may need shortening to fit the mouth that already moved and breaths must go where picture leaves room; the treatment should say which words are fixed and which may be trimmed. Voice first with articulation deliberately pushed is a style choice that should be flagged or someone may correct it later. For testing, the aligned and divergent versions have to be seen and heard together; charts document intention rather than test it. Ask whether the mismatch reads as intended or broken, and ask a few people who said the line whether the answer changes between versions; watch whether Dev's silence plays as a beat or as confusion about the film. If divergence lands as an error rather than a joke, revise the performance proposition rather than assuming cleanup brings the idea back. The article states the variants are written, not recorded; no performer has been asked, no animatic exists, and nothing has been watched or evaluated. The comparison is a plan for one test, not a report, and the vocabulary around lip sync, dubbing, ADR, voice replacement, and localization is unsettled enough that a treatment should check with the post team which practice is meant.

More in Advertising Browse all articles