Skip to content

Put a 3D Product Into Moving Location Footage: Solve the Camera or Use a Planar Mockup?

Advertising

Put a 3D Product Into Moving Location Footage: Solve the Camera or Use a Planar Mockup?

The failure usually shows up late. The first second of the clip looks right, the product sits on the table and the light agrees with it, and then somewhere around two seconds the object should have begun to turn with the camera and hasn't. Nothing in the composite looks broken. It looks like a sticker that has learned to hold still while the room moves around it.

Two routes are available for a treatment test like this, and they are not two versions of the same route. A planar attachment fits a 2D transform to a flat region of the frame and carries your product along with it. A scene-camera solve reconstructs where the camera was, what angle of view it had, and where a cloud of scene points sits in three dimensions, so the product can occupy a place in a modeled room. The first estimates a rectangle; the second estimates a space.

Choose in this order. Ask what has to change about the product as the camera moves. If nothing does — if the placement is genuinely flat content lying in a plane — attach it and move on; a planar mockup is the correct model, not a compromise. If the product's visible form, its contact with a surface, or its depth relative to near objects has to change across the passage, you need a supported reconstruction, and you need to check that the footage can support one before you promise it. And know in advance which frame range would expose the difference, because one convincing frame proves nothing.

Ask what changes with the viewpoint

Make a list of the things this placement has to do across the whole clip you intend to use:

  • Which faces of the product are visible, and how the front face foreshortens as the camera swings
  • Where the base meets the surface, and at what angle that contact is seen
  • The product's screen position relative to near and far scene features
  • What passes in front of it, and what it should pass in front of
  • How its height reads against surrounding objects at both ends of the move

A poster, a wall label, a screen inset, a decal on glass: none of those change with viewpoint, and no amount of tracking sophistication improves them. Pin the flat content to the plane it belongs to.

Now the important part. Suppose you bake a render of a solid product from the camera's viewpoint at one frame and pin that image to the table plane. At the bake frame it is indistinguishable from the solved version. That is the whole difficulty, and it is why the mockup so often survives review. What the plane gives you, frame by frame, is where the plane is. The appliance image rides that transform and inherits the plane's perspective change and nothing else. A solid's near face grows faster than its far face as the camera comes around; a single picture has no near and far and can only be warped as a whole, or sheared, or scaled. You can repaint the difference by hand, or bake a second view and cut to it, and the cut is the tell.

If you find yourself re-rendering for each new viewpoint, you are animating a camera, which is the other route — you have just decided to do it by hand, frame by frame, with no reconstruction underneath it.

Solve from the scene, not from whatever is moving

Everything a camera tracker treats as true rests on one assumption: the world held still while the camera moved. Points belonging to anything that moved on its own — a server crossing, traffic, a curtain, a screen playing video, foliage in wind — describe the camera incorrectly, and the damage is not confined to those points. The solve has to reconcile them with the static ones, so a handful of bad points can tilt a camera path.

Adobe's page on tracking 3D camera movement documents this territory: a tracker that estimates camera motion and scene data, with unwanted-point deletion, a Shot Type setting for a fixed or variable angle of view, and placement of a ground plane and origin. That is procedure documentation. It tells you the controls exist and roughly what they act on. It does not tell you whether your plate is solvable, and it cannot stand in for running the solve on your footage.

Delete points you cannot vouch for, and choose the analysis region for texture rather than for where you want the product. A blank stretch of wall gives the solver almost nothing to hold; a repeating pattern gives it matches it can confuse. Soft focus, motion blur and rolling shutter make matching unreliable, and exposure ramps change the image under the tracker without anything in the room having moved.

Then the geometry of the move itself. Rotation alone, or zoom alone, changes the view without changing the camera's position, and there is nothing to separate near from far. Lateral movement — an arc, a slide, an operator stepping sideways — pulls the table and the back wall apart in the frame, and that separation is what depth is recovered from. The angle-of-view assumption is the same kind of judgment: fixed if the lens didn't zoom and focus stayed put, variable if anything adjusted mid-shot or the shot zooms. If you assume wrong, the solver does not complain. It absorbs the error into a camera path, and the symptom is a camera that moves when it didn't.

A solve also reports a residual — how closely the estimated camera reproduces the points it kept. It is a check on internal consistency, not on the room. Deleting the difficult points usually improves the number without improving the reconstruction, because a smaller, easier set of points is easier to fit. A low residual beside a wrong solve is common enough that it should never be the sentence you send to a client in place of looking at the frames.

Set the ground, the origin and the scale before you build

The solve returns a camera and a point cloud. It does not return a floor. The ground plane and origin come next, and they are decisions rather than discoveries: the origin is where your units begin, and the ground plane is what "on" means for everything you place afterward. A product on a table should be defined relative to a plane you chose, not to wherever it happened to look right in the frame you were staring at.

Scale is the other declaration. A monocular solve recovers shape at an arbitrary size; nothing in the image states that the table is 74 cm high. If a dimension was recorded during the shoot — a tape measure, a slate, a known object, the product's own approved geometry if it appears on camera — use it, and write down which one you used. If nothing was recorded, you can still place the object and the spatial relationship is still meaningful, but the scale is approximate and the deliverable has to say so.

Two disciplines follow from that.

First, if you start nudging the model frame by frame to keep it looking right, you have stopped solving and started attaching. That is a legitimate thing to do under pressure; it just isn't a camera solve, and the notes should not describe it as one.

Second, the model's dimensions come from the approved geometry you were given. If you don't know them and there is no reference in the plate, that is a scale problem, and no amount of tracking work converts it into a measurement of the room.

A shot that has to choose

What follows is a constructed walkthrough, not a report. The plate, the product and the timings are stipulated so the decisions have something concrete to attach to. Nothing here was executed for this article; the native comparison has not been run.

The plate. Four seconds, handheld, fixed focal length as an assumption. The operator steps roughly 60 cm to the right and pans slightly left, netting a change of view of about 25 to 30 degrees; the travel is complete by about 2.2 seconds, after which the operator holds the new position for the rest of the clip, so 2.2 seconds is the widest viewpoint in the shot. A cloth-covered table, 74 cm high, measured at the shoot and written on the slate, with the floor visible beneath it and in frame throughout. A back wall with a poster about 1.5 m behind the table, a window on the left. A server carrying a tray crosses the left third of the frame from about 1.2 to 2.0 seconds. A chair sits in the near foreground; it doesn't move, but because it is close, the camera's own travel sweeps it across the lower frame between roughly 0.6 and 1.2 seconds, so it passes in front of the placement area and clears it again.

The product. An approved countertop appliance, 24 cm tall, 18 cm wide, 15 cm deep, with a front panel and a handle on the left. It goes on the table about 40 cm in from the near edge.

Route one, planar mockup. Track the tablecloth as a plane and pin the appliance to it, baked from the camera's viewpoint at 0.5 seconds. At 0.5 seconds the result is indistinguishable from the solved version. At 2.2 seconds, the end of the travel and the widest viewpoint in the clip, the two are no longer describing the same object: the solved appliance would be showing part of its right side and a foreshortened front panel, and the pinned image shows neither, because neither was in the bake. The chair sweep is roto work in this route — mask the layer and let the chair pass. Note what the crossing does to your evidence: if the layer drifts while it is hidden, you will not see it until the chair clears, and the error accumulated exactly where you couldn't look.

Route two, camera solve. Same plate. Keep the poster, the wall, the table edge and the cloth; delete the server. The two depth layers are the reason this shot can be solved at all — a shot that merely turned on its axis would offer nothing to separate the near cloth from the far wall. Solve with the angle of view declared fixed. Set the ground at the table top, place the origin at a chosen point on the table, and set the model to 24 cm, tying the solver's units to centimeters with the 74 cm measured from the table top down to the floor beneath it. Put the appliance 40 cm in from the near edge and leave it there. The chair sweep is now a depth question with a discoverable answer: the chair is nearer, so the model is occluded and should stay exactly where it was while it is hidden, then reappear in the same relationship.

Where to compare them. Not at 0.5 seconds, and not at the handsomest frame. Compare at the widest viewpoint — here, about 2.2 seconds, where the travel ends — because that is where the two routes are making different claims about the same object.

Then inspect the solved result across the full four seconds:

  • The widest viewpoint: are the visible faces the ones the orbit implies?
  • Contact: does the base meet the cloth at the same place on the cloth at 0.2 seconds and at 3.8 seconds?
  • The crossing: what is true in the frame before the chair covers the placement, and in the first frame after it clears?
  • Drift: pick a cloth feature near the placement and watch the distance between it and the model across the whole clip.
  • Scale: at both ends, does 24 cm still read against the measured 74 cm from the table top to the floor?

A solve that holds at 0.5 seconds and fails at 3.4 seconds has failed. That is the entire argument for looking at the passage instead of the frame.

When the solve doesn't hold

It may not. Too little lateral movement, an analysis region dominated by things that moved, soft or repeating texture under the placement, an angle-of-view change you cannot declare, or a plate that was always going to be stabilized — stabilization changes the camera path, so if the final will be stabilized, the solve has to be done against the stabilized plate or the relationship is a different one.

When that happens, do not go hunting for the frame that flatters the result. The honest options, roughly in order of how much they preserve:

  1. Narrow the claim and the range. Use the interval where the reconstruction does hold, and say in the notes which interval and why the rest was excluded.
  2. Call it a mockup. If what the test actually answers is whether the placement reads — size, position, does the product belong in this frame at all — a labeled planar mockup answers that question cheaply and without pretending to be something else.
  3. Redesign the sample. Shoot a second take built for the solve: more lateral movement, better texture under the placement, focus locked, nobody crossing. A test that demonstrates the relationship under conditions you controlled is a legitimate answer to "does this idea work," provided the difference between the staged shot and the intended shot is written on the label rather than left for the client to discover.
  4. Reduce fidelity. Keep the model in the solved space but deliver stills from the ranges you actually inspected, with the assumptions attached.

Keep the solve that failed. A folder containing an unstable reconstruction and a note explaining why it did not hold is worth more to the next person than a folder containing only the version that worked.

What the test is allowed to claim

The route determines the sentence you can write under the result.

Route What the deliverable can claim What it does not claim
Planar mockup The product reads at this size and position in this composition, and its footprint on the tracked surface holds for the range shown Any change in the product's visible form with the camera; its depth relative to near scene elements; real-world scale
Scene-camera solve A reconstruction carries the model through this passage: view change, contact and the near-chair occlusion hold across the inspected range Measured room geometry; that production will solve this shot quickly, or at all; anything about shots not solved
Redesigned sample The placement idea reads under conditions you controlled The original camera move; that the designed shot is achievable as shot

Whatever the route, the notes travel with the result: the clip range you used, the angle-of-view assumption, which points were removed and why, the origin and ground plane, the scale reference and whether it was measured or approximate, and the frames you did not inspect. Everything else in this article is method. The example above is a construction, and the native comparison of planar and solved placement remains unrun.

End on the range that argues with the placement. Deliver the passage, not the hero frame — the widest viewpoint, the occlusion crossing, the last frame where the depth relationship has to still be true. If the scale is approximate, mark it as approximate next to the object rather than in a footnote. If a reconstruction failed, leave it visible. The client's decision is about feasibility, and the one thing a treatment test must never do is make the shot look easier than it is.

Frequently asked questions

When is a planar attachment the right choice instead of a camera solve?

When the placement is genuinely flat content lying in a plane and nothing about the product has to change with the camera — no visible form change, no changing contact angle, no changing depth relative to near objects. A planar attachment estimates a rectangle; a scene-camera solve estimates a space.

Why can a product pinned to a plane look right in one frame and fail later?

At the bake frame it can be indistinguishable from a solved version. The plane gives only the plane's perspective change. A solid's near face grows faster than its far face as the camera comes around, but a single picture has no near and far and can only be warped, sheared or scaled as a whole. Repainting by hand or cutting to a second bake is the tell.

What makes moving location footage more likely to support a scene-camera solve?

The solve assumes the world held still while the camera moved. It needs lateral movement — an arc, slide or sideways step — so near and far objects separate in frame; rotation or zoom alone leaves nothing to recover depth from. The analysis region needs usable texture, not blank wall or repeating pattern, and moving points should be deleted. The angle-of-view assumption must be set correctly.

What does a low solve residual prove?

Only internal consistency — how closely the estimated camera reproduces the points it kept. It is not proof of the room. Deleting difficult points can improve the number without improving reconstruction, so a low residual beside a wrong solve is common and should not replace looking at the frames.

What should travel with the result when a solve or mockup is delivered?

The clip range used, the angle-of-view assumption, which points were removed and why, the origin and ground plane, the scale reference and whether it was measured or approximate, and the frames not inspected. The route also limits the claim: a planar mockup cannot claim form change with camera or real-world scale; a solve cannot claim measured room geometry or that production will solve the shot quickly or at all.

More in Advertising Browse all articles