Skip to content

Make and Correct a Lip-Sync Test for an Animated-Series Pitch

Television

Make and Correct a Lip-Sync Test for an Animated-Series Pitch

A pitch does not need an animation pipeline. It needs one face, one recorded line, and mouth shapes that arrive on the sound. That is a small job — small enough to finish, inspect, and hand to another person this week.

You already have the two things this test depends on: an original recording you have the right to use, and a puppet whose mouth drawings carry labels the software can read — though a label may point at the wrong drawing. From there the work is three operations. Verify the mapping between each label and its drawing, and correct the tags that point at the wrong artwork. Ask Adobe Character Animator to compute a first pass of mouth cues from the audio. Then correct the timed cues by ear.

That correction step is where the test earns its name. Automatic cues are an interpretation — a guess about which mouth shape belongs at each moment, inferred from the sound and sometimes from a transcript. A guess is a reasonable place to begin and a careless place to stop.

Three objects are involved, and they look like one thing until they disagree. A mouth drawing is an artwork file. A viseme label is the name the software assigns that drawing — its role in the set, such as a closed shape or an open vowel. A cue is an interval on the timeline, with a start and an end, during which a labeled shape is shown. When a mouth reads wrong, you have to decide which of the three is at fault, because the repairs are different and not interchangeable.

Check the mouth set before generating timing

Open the puppet and look at its lip-sync setup before you ask the software to compute anything. Adobe's Character Animator Help page on directly controlled behaviors — the page is marked updated 20 June 2023 — walks through this under its "Working with visemes" section. You are looking for the neutral mouth and the mapped visemes: the drawings that have been tagged for speech.

The check is dull and it saves the whole test. For each label you expect to use, confirm that the drawing attached to it is the one you intended, and retag it if it is not. A closed label should hold a closed drawing. An open-vowel label should hold the mouth you want on a wide sound. If the artwork and the label are mismatched, no amount of timeline editing will fix it — you will be trimming intervals that were never showing the right shape to begin with.

This is the first place the three objects separate. If the wrong mouth appears at the right moment, the fault is in the mapping. If the right mouth appears at the wrong moment, the fault is in the cue timing. Correct artwork mapping and correct cue timing are two different repairs, and confusing them wastes an afternoon.

Building the puppet itself is a separate task with its own decisions about joints, controls, and how much articulation the character actually needs; the companion piece on building a simple digital puppet handles that. Keep this test focused on one recording and the mouth performance it produces.

Generate a baseline from the selected recording

Import the audio into the scene. Character Animator expects a supported format, and the puppet and audio need to sit together on the timeline so the mouth reads against the right sound.

Select the intended puppet track before you compute. Then run the command that generates mouth cues from the scene audio. The exact wording varies between Adobe's Help page and the Learn tutorial — one phrasing runs along the lines of "Compute Lip Sync from Scene Audio," and you should read your installed build's menu rather than trust a memory of the docs. The Help page's section on generating lip-sync data from prerecorded audio describes the operation; the Learn tutorial's transcript notes it around 02:58.

Write down two things the moment the first pass appears: the version of the application you are running and the exact label of the command you used. A cue result is only meaningful next to the build that produced it, and a pitch reviewer who wants to reproduce your take needs both.

Before you touch a single cue, check the range you computed over. It is easy to compute against the wrong audio — a scratch track, a muted take, or a file that overlaps the one you meant to use. Listen to the intended passage and confirm the cues match that sound and not some neighbor.

Then save. Save the generated take as its own file or its own version before any manual editing. The baseline is the evidence that the corrections mattered. If you edit over it and never keep it, your final clip shows a result with no history, and the most interesting part of the test — that you changed something on purpose — is gone.

Correct closures, held sounds and pauses

Now work one mismatch at a time. Play the audio and watch the mouth together, and stop where they visibly disagree. Then look at the cue intervals around that moment, not just the written word underneath them. The transcript or script can mislead you, because punctuation marks intent, not measured time. A period does not tell you how long the mouth stayed closed. Only the audio does.

A plosive is where closed shapes matter. Take a line like "Put it back. Ooh… better." — a proposed line for this test, not a recording made here. The p in "Put" and the b in "back" are closures: the lips meet, then release into the vowel. If the closed shape flashes for an instant and pops open too early, the mouth looks like it is swallowing the word. If it lingers, the character looks like it is holding a sound it never made. Trim or extend the interval until the closure sits where you actually hear it.

A held sound is where interval length carries the performance. "Ooh…" is a rounded vowel you intend to hold. That means one shape persisting across a long interval, not a flicker of the correct shape between two wrong ones. Extend the appropriate vowel cue through the held portion so the mouth stays open for the whole sound. This is a timing repair, not a new drawing.

A pause is where the documented behavior trips people up. Adobe's Help page states that deleting a viseme extends the preceding interval — the earlier shape simply continues into the gap. That is exactly what you do not want during silence, because the character will sit there mid-speech with no sound to justify it.

For an actual pause, replace the cue rather than removing it. Select the cue that occupies the gap, then set that cue's viseme to Silence. The Learn tutorial covers this around 02:20 in its cue-editing section. With the gap still selected, confirm that the shape now shown through it is the puppet's neutral mouth — the same neutral shape you identified when you checked the mouth set. Do not assume Delete produced silence, because it did not: deleting extends the shape before the gap, while the Silence viseme holds the neutral mouth for the interval you gave it.

Here is the failure worth naming, so you can avoid it. You hear a pause, you want the mouth to stop, and you delete the cue sitting under that gap. The software extends the shape before it, and now the character holds "Ooh" straight through the rest of the line. You have made the timing worse and harder to diagnose, because the timeline no longer shows the problem — it shows a long vowel cue that looks deliberate. The repair is to replace, not delete, and the difference lives in the documentation.

Work the three corrections in order — closure, held vowel, pause — and playback after each one. Corrections interact. A closure you move can push the vowel that follows, and a Silence you insert can shorten a hold you just fixed.

Separate the mouth correction from other performance

A fair test changes one thing at a time. If head motion, expression, or unrelated behavior recordings are active while you judge a closure, you cannot tell whether the mouth improved or the extra acting simply distracted from it.

For the comparison, keep those other tracks stable or disable them. Play the baseline and the corrected passage under the same conditions — same audio, same playback speed, same screen. Then look at the mouth alone. Ask whether the closure you targeted now lands on the right sound, or whether you have hidden a new mismatch under a head turn.

Automatic detection stays a first pass. The Help page notes that transcript-assisted processing is best suited to English and ASCII text, with corrections expected. Read that as a scope limit, not a dare. Do not carry an English-oriented assumption into a test in another language, and do not prescribe mouth cues from spelling alone — "better" in American English can pass through a soft flap rather than a crisp t, and the sound, not the letters, is what the audience hears. A single-language result from a single recording is a single-language result.

Save the test so another person can inspect it

A pitch test that only exists as memory is not a pitch test. Package it so someone else can open the work and reach the same conclusion you did.

Keep four things together. The original audio, untouched. The puppet with its mouth mapping visible. The editable cue track, with the corrections still live. And a labeled clip of the provisional performance. The label matters: mark it clearly as a provisional test, so no one mistakes it for finished character animation or for the approved voice of the series.

Show the baseline and the corrected passage as a pair, at the same playback conditions. That pairing is the whole argument — here is what the software produced, here is what you deliberately changed. If you cannot show the before, the after is just a mouth moving.

Then state the limits out loud, in the same breath as the result. This test examines one recorded line on one face. It does not establish finished character acting, the full series pipeline, the cost of doing this across an episode, or whether the production can sustain it. It shows that you can take a recording, generate a first pass, and correct the mouth where it fails — and that is a specific, checkable claim.

End on the artifacts: the paired clips, the editable take, and the three observed corrections. Report what happened on this recording. A reviewer who sees the closure you tightened, the vowel you held, and the silence you inserted will trust the rest of your pitch more than any claim that the character animation is solved.

Frequently asked questions

What three elements can disagree when a mouth reads wrong?

A mouth drawing (the artwork), a viseme label (the name that assigns its role, such as closed or open vowel), and a cue (the timeline interval showing a labeled shape). Wrong mouth at the right moment points to mapping; right mouth at the wrong moment points to cue timing.

What should be checked before generating lip-sync cues?

Verify each label is mapped to the intended drawing and retag mismatches. A closed label should hold a closed drawing; an open-vowel label should hold the intended wide mouth. Timeline editing cannot fix a bad artwork-label mapping.

How do you correct a pause?

Replace the cue in the gap with the Silence viseme rather than deleting it. Deleting extends the preceding shape into the pause, so the character may appear to hold a sound through silence. Confirm the neutral mouth shows through the gap.

How are closures and held sounds corrected?

For plosives such as p and b, trim or extend the closed-shape interval until the closure lands where it is heard. For a held rounded vowel like 'Ooh,' extend the appropriate vowel cue through the held portion rather than leaving a flicker.

What should be saved for a pitch-ready test?

Keep the original audio untouched, the puppet with visible mouth mapping, the editable cue track with corrections, and a clearly labeled provisional clip. Show the baseline and corrected passage together under the same conditions, and state limits such as one recorded line on one face.

More in Television Browse all articles