What Must Survive When a Human Performance Becomes an Animated Character?
What Must Survive When a Human Performance Becomes an Animated Character?
Suppose a treatment hands the animation team a clip of a performer and one line of direction: match the performance. The performer turns toward someone, greets them, and moves a hand out of sight. The animator watches the clip until it stops yielding anything and copies the shapes. Often the shapes are fine. What has gone missing is something else, and it went missing before the first keyframe.
A recorded performance holds two things at once: an action, and one body's way of performing that action. Only the second is on the surface, where it can be measured and copied. The first has to be inferred, and the inference is the actual animation work. If the brief never states it, the animator supplies one silently — usually a plausible one, sometimes the wrong one, and the wrongness is nearly impossible to argue about later because the footage really does look like that.
So the useful answer is this: preserve what the performer is trying to do to somebody, and treat the original joint positions as one available expression of it. Everything below is a way of doing that on purpose rather than by accident. To keep it concrete I'll build one invented moment, a person and a creature, and carry both through the same six seconds. No footage, capture, rig, or animation underlies the comparison; it is worked on paper.
Describe the action beneath the pose
Start with something the performer is doing to someone or something, because that is the only kind of instruction an animator can act.
Hide the key is not playable. It describes an outcome and a verdict, and there is no way to mean it. Keep the visitor from finding out about the surprise is playable. It has a target, a stake, and a reason to keep moving as the visitor moves. The concealment of the key is then a subordinate move inside a larger intention, not the intention itself, and that distinction controls every later decision about which arm does what.
Now the invented moment. A performer stands at a table. A small brass key lies on the near-left of the surface. A stool sits to the performer's right. A visitor will enter at the far side of the table, facing the performer, and the key is in plain view from that entry position. The key is a surprise for the visitor, so it must not be seen. The performer's intention, stated as a direction an animator could play: give this visitor a warm welcome, and let them find out about the surprise later.
Separate that intention from the means. Holding a key in a closed hand, carrying it to the small of the back, shifting weight onto the right foot, opening the other palm toward the stool — these are one body's solutions. Each one is replaceable. What is not replaceable is what they are for.
Mark where the intention changes, because the change is usually the beat. Before the visitor appears, the performer is managing a private situation. The moment the visitor enters frame, the intention turns outward, and everything downstream reorganises: gaze first, then the torso, then the hands, then the feet. An animator who knows the turn is coming can build the earlier seconds to make it land.
And mark the response the action causes, since a response is what turns a pose into a scene. Here the visitor walks toward the stool. That is the evidence the welcome worked. A striking silhouette proves none of this. A body can be beautifully shaped and still be doing nothing to anybody, and that failure is invisible in a still frame and obvious in motion.
Find the signals the new body can carry
List the signals in the reference moment separately: gaze, weight shift, pause, contact, release. Then map each against the anatomy you actually have. The mapping produces three outcomes, and naming which outcome applies to each signal is most of the transfer.
In the human version the signals run like this. The gaze drops to the key, then lifts to the visitor's face and stays there. Weight moves onto the right foot as the torso squares to the visitor. The left hand closes on the key and travels behind the hip. There is a short pause before the right hand rises. Between the two arms there is an asymmetry the audience can read: the right arm is doing something for the visitor, the left arm is doing something for the performer.
Which of those transfer to a different body? Gaze transfers if the character has eyes, and here the gaze is the load-bearing signal, because it is what tells the audience the key exists and matters. Pause transfers, because timing is a property of the edit and the performance rather than of the skeleton. Weight shift transfers only partly: a character with feet can carry it almost literally, a character without them needs some other way to show the body committing to a direction.
Contact does not transfer, and this is the important one. The human concealment works because a closed hand holds the key. Holding is a state, not a gesture. Once the hand closes, the concealment is maintained with no further movement at all, for as long as the shot runs, at any point in the room. A character without hands cannot hold. If the key stays on the table, the concealment has to be maintained by position — the character's mass has to sit on the line between the visitor's eye and the object, and it has to stay there.
That single difference drives everything. A hand can conceal and then do nothing. A body can only conceal by staying put, and staying put conflicts with offering. So for a character without hands, the interesting question is not which gesture replaces the hand, but what the offering has to become so that the concealment can survive it. A creature that leans across the table, or turns to lead the way, or moves toward the stool has just uncovered the key. It has no way to conceal except by remaining in place, which means the welcome has to be made from where it stands.
Which is why the tempting substitution — copy the human hand onto whatever limb is nearest — fails for a reason that has nothing to do with looks. Suppose the proposed character has a short stubby arm. Making it perform a hand-cover gesture gives you a shape that resembles a hand covering something. It does not put the stubby arm's mass into the sightline between the visitor and the key, so the visitor could still see the key, and the character's action no longer does anything. The instruction "hide the object" was translated into the instruction "make the shape of hiding," and the second is not the first.
Change the body position rather than copying the limb.
Compare imitation with an authored performance
Two interpretations of the same beat are worth writing down on the same timeline, because the comparison is where the decision becomes arguable instead of aesthetic. Take one beat of the invented moment, six seconds long, in both bodies.
Human version, close to the reference.
- 0:00–0:01 — Alone at the table, the key held in the left hand at waist height. Gaze down on it.
- 0:01–0:02 — The visitor enters at the far side. Gaze lifts to the visitor's face. Head turns, shoulders follow. The left hand closes and travels to the small of the back. Weight onto the right foot.
- 0:02–0:04 — The right hand rises and opens toward the stool. One small nod. Gaze holds the visitor's face.
- 0:04–0:05 — The open gesture holds. The left shoulder settles down and back.
- 0:05–0:06 — The visitor steps toward the stool. The right hand lowers to the table's edge.
Where does concealment become readable? Not at the hand. A hand behind a hip is just a hand behind a hip. It becomes readable at 0:01–0:02, at the moment the gaze leaves the key — because the audience already watched the performer look down at something and hold it. Cut the first second and the whole beat collapses: an audience sees a person turning, greeting, and putting an empty hand behind their back. There is nothing to conceal and therefore no scene.
Notice also what the greeting is doing. Between 0:02 and 0:04, if the visitor's gaze drops from the performer's face to the arms, the asymmetry is plainly visible. The performer's held gaze, the nod, and the open palm are keeping the visitor's attention up and to the right, away from the near-left of the table. The offer is politeness and misdirection in the same gesture, and the second job is what makes the beat work.
Creature version, authored for the body.
The invented character: rounded, handless, about the height of the tabletop, with two forward-facing eyes with lids, a mouth line, and a broad base. It can turn in place, tip its upper form toward a direction, and rise or settle. The key lies on the near-left; the creature sits beside it.
- 0:00–0:01 — Beside the key, gaze down on it.
- 0:01–0:02 — The visitor enters. Gaze lifts. The body settles low and shifts slightly left, deepening the occlusion of the key from the entry sightline.
- 0:02–0:04 — The upper form tips toward the stool on the right. Mouth line opens. Gaze holds the visitor's face. The base does not move.
- 0:04–0:05 — The tip is sustained as a hold.
- 0:05–0:06 — The visitor steps toward the stool. The creature stays put.
The step at 0:05–0:06 is where this version's geometry has to be stated rather than assumed. The visitor's route runs along the far side of the table toward the stool, which sits on the performer's right — away from the key's near-left corner. Each step along that route carries the visitor's eye further from the key's side of the table, so the sightline to the key crosses at a steeper, more oblique angle, and the creature's settled mass, beside the key, stays across it. The concealment holds because the visitor walks away from the key's side rather than toward it. A visitor who circled toward the near-left would put the creature's fixed position behind the sightline instead of across it, and a body that cannot hold an object would then have to choose between tracking the visitor and making the welcome from where it stands. The human version never had to face this: a closed hand holds the key wherever the performer stands, so the visitor can walk anywhere.
The problem sits at 0:01–0:02. In the human version, concealment had a channel of its own: a second arm, moving separately from the arm doing the welcome, which left the pause before the offering free to serve the offering alone. In the creature version, the concealment and the turn-to-greet are one movement. There is no separate signal for I do not want you to see this. So the authored version has to give one pause two jobs — a half-beat hold before the settle, so the settle has an onset the audience can notice, or a downward glance that does not quite release the visitor's face. Pause transfers; the reference already contains one. What is authored here is where it falls and what it has to carry.
And it costs, though not in a new element. The reference spends its six seconds with a pause already inside them, and that pause was affordable because the arm was handling the concealment. Move the pause onto the concealment and it has to be long enough to read as concealment, which either shortens the offering movement at 0:02–0:04 or pushes the beat past 0:06. There is no free time. The creator who wants both a readable concealment and a generous invitation from a single mass has to decide which one thins out.
The close-imitation creature version is easy to build and easy to describe: dip the left flank down toward the key and lift the right flank, holding the pose, mirroring the human hand that went behind the hip. What an audience sees is a sway. No part of the creature holds the key, the key is still on the table, and the flank dip changes nothing about whether the visitor can see it. The pose was copied; the action it was performing was not.
Do not assume the authored version wins on that basis, though. It buys a concealment that reads with an invitation that is smaller, a hold that is harder to animate than a gesture, a concealment that survives only while the visitor's route carries them away from the key's side of the table, and a beat that depends on the visitor's attention staying where the creature put it. It is entirely possible that the single-mass version, if it works, reads as one intention rather than two, and is stronger than the human division of labour — or that it reads as a shy object doing nothing at all. Nothing on the page decides it. The comparison tells you what to test, not what the answer is.
Show the treatment's actual level of proof
Treatments get argued about because four different kinds of material get labelled with the same word: the reference.
A recorded reference shows what a person did in one take. It can establish that a specific action is legible on camera, at that pace, from that angle. It cannot show what a different body will do with it.
Motion capture is a measurement of a body's movement, and a detailed one. It is material, and good material, in the way that a photographed location is material. It does not decide which part of the movement carries the acting, and it does not solve a character whose proportions or limbs do not match the source.
Rough animation is a proposal: an animator's claim about what the action should become in the new body. It is evidence that somebody made a decision. It is not evidence the decision reads.
A finished test — a shot-length pass at final quality — supports a narrow and useful claim: this character can be animated reading this way, at this quality, in this shot. Even then it is one reading, and it has not been shown to an audience unless it has actually been shown to one.
On structure, there is some public material. Epic's Unreal Engine documentation on animation retargeting, whose introduction and "Why Use Retargeting?" section were checked on 8 September 2026, describes reusing animation across differently proportioned characters and explains why those proportions require attention during a transfer. That is a real and relevant point, and it is narrowly about structure. Nothing there establishes that a performer's intention, expression, or identity survives the transfer, and only the page text was inspected — not the illustrated animation, a rig, or an executed capture transfer. Retargeting can be technically clean and dramatically empty. A rig can reproduce every joint angle of the reference and produce a character that is doing nothing to anybody.
What is missing here is a first-person account from an animator describing an actual adaptation decision — what was in the reference, what the new body could not carry, and what was invented instead. That is the evidence this argument most needs, and it is not part of what this piece can offer. The paired comparison above is worked on paper; no recording, capture, animation, rig test, or audience assessment has been performed.
The transfer brief
This is the shape of the thing to hand an animator, and it is deliberately short.
Intention to preserve. Give this visitor a warm welcome; keep the surprise hidden until the visitor finds it. The intention is directed at the visitor, not at the object.
Signals to adapt. Gaze down on the object, then released to the visitor's face — keep this, it is what makes the concealment readable at all. The pause before the offering move — keep it, and consider lengthening it if the concealment needs an onset of its own. Weight shift and the closed-hand contact — reassign or drop. Do not build a limb that makes the shape of a hand covering something.
Proposed physical interpretation. The character's mass stays across the visitor's sightline to the object, for the length of the shot. Check that against the visitor's route and not only their entry position: the concealment holds while the route carries the visitor away from the object's side of the table. The welcome is made from the upper body and the eyes, with the smallest possible change of position. The offering move points the visitor's attention away from the object's side of the frame.
Test still needed. Animate both versions of this beat — a close imitation of the reference and the authored reinterpretation — in the actual rig, at shot length, and show them to someone who has not read this brief. Ask that person what the character is trying to do, and to whom. If they describe a mood instead of an action, the concealment is not readable yet.
Frequently asked questions
What is the central thing to preserve when a human performance becomes an animated character?
Preserve what the performer is trying to do to somebody, and treat the original joint positions as one available expression of that intention. The action has to be stated as a playable direction, not an outcome or verdict.
Why is 'hide the key' not a playable direction, while 'keep the visitor from finding out about the surprise' is?
'Hide the key' describes an outcome and a verdict, with no way to mean it. 'Keep the visitor from finding out about the surprise' has a target, a stake, and a reason to keep moving as the visitor moves; concealment of the key becomes a subordinate move inside a larger intention.
Which performance signals transfer to a different body, and which do not?
Gaze transfers if the character has eyes, and here it is load-bearing because it tells the audience the key exists. Pause transfers. Weight shift transfers only partly, depending on feet. Contact does not transfer: holding is a state, not a gesture. A character without hands must conceal by keeping its mass on the sightline between the visitor and the object.
Why does copying a hand-cover gesture onto a stubby arm fail?
It produces a shape that resembles a hand covering something, but it does not put the stubby arm's mass into the sightline between the visitor and the key. The visitor could still see the key, and the character's action no longer does anything. Change the body position rather than copying the limb.
What kind of proof do reference, motion capture, rough animation, and a finished test each provide, and what test remains?
A recorded reference shows what a person did in one take. Motion capture measures a body's movement but does not decide which part carries the acting or solve a mismatched character. Rough animation is a proposal. A finished test supports a narrow claim: this character can be animated reading this way, at this quality, in this shot; even then it is one reading and not audience-tested unless shown. The remaining test is to animate both the close imitation and the authored reinterpretation in the actual rig at shot length, show them to someone who has not read the brief, and ask what the character is trying to do and to whom. If they describe a mood instead of an action, the concealment is not readable yet.