Write a Virtual-Production Treatment Around What Must Be Built Before the Shoot
Write a Virtual-Production Treatment Around What Must Be Built Before the Shoot
There is a version of the virtual-production pitch in which the background is a decision you make later. The deck has a hero frame. The volume is booked. The environment is treated as something that arrives when the director does, the way a location arrives when the van does.
An in-camera environment does not work like that. The moment a background is photographed instead of composited, its contents become part of the shoot day rather than part of the weeks after it. Geometry, movement, light and — most easily forgotten — which parts of it the lens will actually see all have to be settled before anyone rolls. The treatment is where that gets said, or where it quietly gets skipped.
What the treatment should do, then, is describe the proposed shot as a relationship between three things: what the performer does, where the camera looks, and what has been prepared behind them. Then it should say what has to exist before those three can be tried together.
The example below is invented to show the method. Nothing in it has been built, tested or captured. Keep it in mind as a shape rather than a case study.
Describe the action across foreground and environment
Begin with the action, not the wall. A performer sits at a table on a train, picks up a travel mug and holds it. The camera needs a close view of the product — lid, label, the surface of the mug filling most of the frame — and a short move that changes the relationship to the platform outside.
Now sort the elements, because they are not all the same kind of thing. The performer, the hand, the mug, the table, the seat and the carriage window frame are physical. They exist on the day, they are lit on the day, and they produce the contact and weight the camera is there to catch. The station beyond the glass is not. It is displayed, and its whole job is to be somewhere that a train might actually be.
That sorting immediately produces a design question that a background image cannot answer. The mug sits still on the table. If the train is moving, the only evidence of that motion in the frame is outside the glass. So the environment is not scenery standing behind the action; it is carrying the premise. If the platform's motion is wrong — too fast, too slow, or entering the frame at the wrong moment relative to the performer's glance — the shot reads as a product on a table in front of a screen.
This is the point at which a treatment earns its keep, because it forces someone to name the background changes that matter. Not "station exterior." Rather: the platform edge and its markings; the vertical elements — pillars, signage, lamp posts — that give the passing motion something to be measured against; the deceleration as the train comes in; the specific moment a particular structure crosses the window. A treatment that lists only the station's appearance has specified about half of what the environment has to do.
Specify a range of useful viewpoints
Next, describe the angles the scene needs, and describe them as a range rather than as one picture.
A single hero frame is a forgiving artifact. It can be drawn, painted, rendered and approved while hiding two expensive problems: environment detail that does not exist outside the one angle the frame was made from, and a mismatch between where the physical foreground sits and where the virtual background expects it to sit. If the camera never moves and never reframes, neither problem surfaces. The treatment should not arrange that by accident.
So say what has to hold: the close product view, where the mug is large and the window is at the edge or out of frame entirely; and the move, which ends with the platform visible behind the performer. That move is the reason the environment needs depth. Over the course of the travel, the carriage window frame — physical, close to the lens — and the platform — virtual, further away — have to shift against each other at a rate that looks like a real distance between them. A flat image can slide as a plane. It cannot make the platform's own contents pass the pillars at their own correct rates, because that relative motion lives in the environment's geometry, not in its picture.
The range also has to be stated in operational terms and left genuinely open inside them. Something like: the camera may be at these marks, at roughly this height, with the mug occupying between this much and that much of the frame. That is a range a collaborator can price and build against, and it still leaves the operator room to find the move on the day. Locking every position is the other failure — it buys certainty the scene may not need and spends preparation on framings nobody will use.
One more discipline. Flag what has only been illustrated. If a framing exists as a storyboard panel and has never been explored in an environment built to receive it, the treatment should say so, in those words, rather than letting a drawing do the work of a test.
Separate illustration, environment and stage readiness
Three different artifacts get conflated constantly, and each one is evidence of a different thing. Keeping them apart is most of what "what must be built before the shoot" means.
An illustration communicates intention. It can be a frame, an animatic, a painted concept. It settles what the shot is supposed to feel like. It establishes nothing about whether the environment behind the performer exists, moves correctly, or holds up from the viewpoints in the range.
A prepared environment is material someone can actually explore in the engine: the platform modeled to the needed depth, its motion authored, its scale and position matched to the physical set. This is a real advance on the illustration, and it still only supports exploration. It tells you the background exists. It does not tell you the camera tracks against it, or that the light on the mug agrees with the light supposedly arriving through the window.
A joint camera-and-stage test is the only one of the three that says anything about capture. Camera, lens, tracking, stage lighting and the physical foreground in one place, examining whether the recorded image holds.
It is worth being precise about why the third is separate, because the dependencies are documented rather than folklore. Epic's In-Camera VFX Overview (checked 18 September 2026) connects LED-displayed virtual environments to tracked physical cameras, and describes a calibration workflow that aligns virtual and physical camera information, alongside lighting dependencies. The same overview separates the region the physical camera frames from a surrounding region the system renders beyond it — which is one reason the environment cannot consist only of the rectangle visible in the approved hero frame. Material just outside that rectangle is not automatically free.
Two limits belong here, stated plainly. That documentation is product documentation. It establishes that these dependencies exist, not that any particular stage can deliver this shot, and it says nothing about cost, schedule or whether the proposal is any good. And the documentation was not accompanied by any live inspection of tracking, lens, LED volume or render for this piece. The useful move is to convert it into questions for the people who would actually build the thing: is the camera tracked here, and by what procedure is it aligned with the virtual one? How is stage lighting matched to the environment's light? What happens to the environment just outside the framed area?
Move consequential preparation ahead of the joint test
Some decisions cannot wait for the stage day, because they change what the stage day has to prove. The treatment should route them early.
For the mug shot, the consequential items are few enough to list. The geometry: where the platform edge sits relative to the actual window frame, at the start and end of the move. The motion: the deceleration profile and the moment the platform structure crosses the frame against the performer's glance. The lighting relationship: what the supposedly outdoor light does to the performer's face and to a mug that is likely to have a glossy or reflective finish. And the contact: the hand taking the mug off a table that does not move while the world outside does.
Each of those belongs to a named owner before capture — the person building the environment, the cinematographer, the person working the lighting, whoever runs the stage system — and each one becomes a question the joint test has to answer. Does the mug's surface pick up the display in a way the product close view can tolerate? Does the platform edge line up with the physical window frame across the whole travel, or only at the ends? Does the platform's stop read at the pace of the performance? Each has a yes, a no and a fix, and the treatment is more useful if it says which one is being asked for rather than assuming the answer.
Then the trade-off, honestly stated. A broader prepared viewpoint range buys freedom: the operator can find the move, the shot can survive a later directorial change, and the environment does not have to be re-authored mid-shoot. It costs preparation, because the platform has to hold up from more positions and the surrounding region has to exist. A narrower range costs less and gives less. If the move cannot be supported, there are real options: shorten the travel so the platform is seen briefly and at a shallow angle; keep the wider framing but lock it off, so the platform still has to move but only from one viewpoint; or split the shot, letting the close product view never show the window at all and taking the wider platform moment separately. Any of those is a legitimate answer. None of them is free, and none of them is a reason to leave the decision to the morning of the shoot. Nor does any of it remove work after capture — photographing the background in camera changes when the work happens, not whether there is any.
The sequence, and the question each step answers
What the treatment is really proposing is an order, with a question attached to each artifact.
Write the action and the frames first, and answer: what does the audience see, in what order, and what in the frame proves the train is moving?
Then the viewpoint range, and answer: from which positions, heights and framings does the environment have to hold?
Then the environment build to that range, and answer: does the platform exist, at the right scale, in the right place, moving at the right rate, from all of those positions?
Then the joint test of camera, light and foreground, and answer: do the three agree in a recorded image?
Only after the last one can anyone say the shot is ready — and if it has not happened, the treatment should say that too, in the same plain sentences it uses for everything else. The uncertainty is not a weakness in the document. It is the part that tells a producer what to fund and a stage what to prepare, which is a great deal more useful than one immaculate picture that answers nothing.
Frequently asked questions
Why can't an in-camera virtual background be treated as a decision for later?
Once a background is photographed instead of composited, its contents become part of the shoot day rather than the weeks after it. Geometry, movement, light, and which parts the lens will actually see have to be settled before anyone rolls. The treatment should describe the shot as a relationship among what the performer does, where the camera looks, and what has been prepared behind them, then say what must exist before those three can be tried together.
What is the problem with approving a single hero frame?
A hero frame can hide environment detail that does not exist outside that one angle and a mismatch between the physical foreground and the virtual background. If the camera never moves or reframes, neither problem surfaces, but the treatment should not arrange that by accident. It should state a range in operational terms: where the camera may be, at roughly what height, and how much of the frame the subject may occupy.
What is the difference between an illustration, a prepared environment, and a joint camera-and-stage test?
An illustration communicates intention and settles what the shot is supposed to feel like; it establishes nothing about whether the environment exists, moves correctly, or holds from the needed viewpoints. A prepared environment is material someone can explore in the engine, with the platform modeled to depth, motion authored, and scale and position matched to the physical set; it supports exploration but not capture. A joint camera-and-stage test is the only one that says something about capture, because camera, lens, tracking, stage lighting, and physical foreground are examined together in a recorded image.
Which documented dependencies make in-camera VFX more than a background image?
Epic's In-Camera VFX Overview connects LED-displayed virtual environments to tracked physical cameras, describes a calibration workflow that aligns virtual and physical camera information, and notes lighting dependencies. It also separates the region the physical camera frames from a surrounding region the system renders beyond it, which is one reason the environment cannot consist only of the approved hero rectangle. The limits are that this is product documentation: it establishes dependencies, not that a particular stage can deliver the shot, and it says nothing about cost, schedule, or quality.
Which decisions must move ahead of the joint test, and what trade-offs belong in the treatment?
Geometry, motion, the lighting relationship, and contact should be routed early and given named owners before capture. They become questions the joint test answers: whether the mug's surface picks up the display, whether the platform edge lines up with the physical window frame across the travel, and whether the platform's stop reads at the pace of the performance. A broader prepared viewpoint range buys freedom but costs preparation; a narrower range costs less and gives less. If a move cannot be supported, shortening the travel, locking off a wider framing, or splitting the shot are real options, but none is free and none should be left to the morning of the shoot.