Build a Two-Eye Stereo Mockup—or Explain the Depth in a Flat Test?
Build a Two-Eye Stereo Mockup—or Explain the Depth in a Flat Test?
A stereoscopic mockup is two things at once: a pair of views that differ in exactly one controlled way, and a route by which a named person can look at that pair as a pair. The rig is the easy half. The route is the half that gets skipped, and it is the half that should decide whether you build the rig at all.
So the rule is blunt. Build the two-eye mockup when you know what the recipient will open it in and how they will look at it. When that is unknown, write the flat breakdown, mark it flat on the artifact itself, and leave the binocular question open rather than closing it with a file nobody can use.
And keep this separate from the question of whether the frame looks deep. Shallow focus, a wide lens, a strong vanishing perspective, three layers staggered at tidy intervals—all of that reads as depth in one image, and none of it contains information about a second view. A reviewer who says the frame "feels deep" has told you about the composition. They have not told you anything about the pair.
Ask what the reviewer can actually open
Three questions, all answered before the scene is built, because the answer determines the export arrangement and the arrangement constrains the scene you can shoot in.
What can they open? The format families are not interchangeable: an anaglyph survives any color screen but spends color to do it; a full side-by-side pair needs a viewer or a display already in the matching mode, and if you pack both images into one ordinary frame you have halved the horizontal pixels for each eye; over-under does the same thing vertically; a headset-targeted file needs the headset and an app willing to load it; an autostereoscopic panel needs that panel.
On what display, in what mode? A stereo-capable monitor in the wrong mode is a flat monitor. So is a shared screen on a review call—treat a video conference as a flat review unless someone has specifically checked that the stereo arrangement survives the pipeline, because usually it does not.
In what room? One person at a desk with a headset is a different review from six people around a conference table watching one screen. Write down which one you're actually getting.
Now name the beat that needs two eyes. Not the whole spot—the beat. In the scene discussed below, it is the moment the near card crosses in front of the product. In a flat frame that moment is occlusion, and occlusion is a perfectly good thing to show. In the pair, each eye sees a different width of product past the card's edge, and the card's position against the backdrop shifts by a measurable amount from one eye to the other. That difference is what the pitch is claiming. If the beat survives intact when both eyes receive the same image, stereo isn't buying you anything, and the earlier decision—whether this commercial needs binocular depth at all—has not been settled yet.
Where the receiving condition is unknown, do not substitute an impressive export for an answer. A beautiful pair with no stated route is a more expensive way of not answering the question.
One scene, two eyes, and a scale you write down
Build one master scene and derive both views from it. The scene can be small: a near card, a product-shaped placeholder in the middle, a backdrop plane. Use an extruded box or a simple shape for the product, not a finished asset—the test is about the pair, not the model, and a placeholder keeps the file light enough to iterate.
Fix the distances and say what the unit means. For this walkthrough the camera sits still, the card is 40 units away, the placeholder is 120, the backdrop is 400, and the eye separation for the first pass is 1 unit. Convergence is by horizontal shift, no rotation—the two eyes stay parallel and one is offset sideways from the other—and the 120-unit placeholder is the nominated reference plane. Keeping the shift lateral is what holds every offset in the table below on a single horizontal axis. The unit is arbitrary; the arithmetic is not. If you don't define it, none of the numbers below mean anything and neither will your notes to anyone else.
Those numbers are stipulations. Treat the whole scene as a construction on paper, not a record of a session—nothing described here has been rendered, exported, or viewed on a stereo display, and the checks below are written as procedures rather than observations.
Adobe's documentation for stereoscopic 3D in After Effects describes a rig that builds left and right compositions from a scene and a combined format for viewing them, along with separate controls for the pair and an arrangement for inspecting the two views (help page; the page was noted as checked on 2026-09-09). Control names move between versions, so confirm the current wording before relying on any label. The reason to work this way is identity: both eyes must contain the same objects, in the same positions, at the same time, with the same materials. The only reliable way to guarantee that is to keep one place where the scene lives and regenerate both eyes from it. The moment you hand-edit the left-eye composition for a look change, the pair has stopped being a pair and become two similar shots.
Label eye identity in the filenames and, if you can, in a slate frame at the head of the clip. The label is the only thing that survives the trip through email, shared drives and whatever player the recipient has. An unlabeled pair is a pair nobody can audit, including you three weeks later.
Two things this method is not. It is not stereo reconstruction from an unknown photograph or a supplied hero render: you cannot derive a second view from one view, you can only invent one, which is a different claim with a different burden of proof. And it is not a license to render the eyes separately and hope. Render both over the same frame range. A pair that is one frame out of step produces a horizontal shimmer that has nothing to do with depth and everything to do with your render settings.
Separation and convergence are two different knobs
With parallel cameras, the horizontal offset between the two images for a point at distance z is the focal length in pixels times the separation, divided by z. That single relationship explains most of what goes wrong.
Separation—the distance between the two eye positions—sets the size of every offset in the scene. Double it and every offset doubles. Nothing else changes.
Convergence decides where the zero-offset plane sits: the plane that lands on the screen. Everything nearer than that plane is offset one way, everything farther is offset the other way. This is the reference the audience's eyes actually use, and it is the plane you nominate, not one you inherit.
Parallel cameras put zero offset at infinity, so the reference plane is always a deliberate choice, made by rotating the cameras toward a common point or by shifting them without rotation. Those two are different geometries—rotation introduces a vertical component that a horizontal-offset model doesn't account for. Pick one and know which you picked. The walkthrough above takes the shift, which is why the offsets it predicts stay on one axis and the swap test below can be read as a single horizontal reversal.
Now do the arithmetic before you build. With a stipulated focal length of 1000 pixels and a separation of 1 unit:
| Layer | Distance from camera | Offset between eyes | Offset relative to the reference plane |
|---|---|---|---|
| Near card | 40 units | 25.0 px | +16.7 px (near side) |
| Product placeholder | 120 units | 8.3 px | 0 (the reference plane) |
| Far plane | 400 units | 2.5 px | −5.8 px (far side) |
Read the last column as a prediction you can check against the render. The card sits at a third of the reference distance, so it carries about three times the reference plane's own offset; the far plane at 3.3 times the distance carries about 0.3 times. Against the reference, the card's excursion is roughly 2.9 times the far plane's, on the opposite side. These are consequences of the 1/z relationship, not preferences, and you can verify each one with a ruler tool on the two images. If the rendered pair disagrees with the table, either the reference plane isn't where you think it is or something is wrong with the pair—both worth discovering before a client does.
Then inspect in two passes. Look at each eye on its own, as a flat image, checking identity and composition: same objects, same frame, same materials. Only after both eyes pass that test look at the combined arrangement and check alignment. The documented rig provides a comparison view for exactly this, and using it before the single-eye check usually means you're debugging two problems at once.
Record what you used in one line you can paste into the review email: scene version, separation, convergence reference and the mechanism that set it—shift or rotation—export arrangement, viewing route. That line is a record of a condition this test ran under. It is not a claim that anything is safe to watch, and no number in it should be read as a house standard.
Break it on purpose
Before you trust the pair, make a deliberately swapped-eye version and keep the intact one beside it. Duplicate the export, swap which eye carries which view, and name the fault file so plainly that it cannot be mistaken for a deliverable later.
The diagnostic has a useful asymmetry. The reference layer is a terrible tell, because its offset is zero in both versions—the swap doesn't move it. The near card is the tell. In the corrected pair, the card's offset is large and falls on the near side of the reference plane. In the swapped pair the magnitude is identical and the side is reversed, which means everything closer than the reference plane now reads as behind it and everything behind reads as in front. Measure the offset in the two images rather than describing what it looks like; the geometry is checkable without fusing anything, and what any individual viewer reports when they do look is a perception question that belongs to whoever runs a proper viewing test.
A second diagnostic worth building once: a layer present in only one eye. It happens for structural reasons—someone adds a caption to one composition after the pair was generated, a copy-paste lands on one side, or a setting pins the layer to a single eye. This is a content mismatch, not a depth decision, and it is the kind of fault that gets mistaken for one. The combined view has no second view of that layer to compare against, so there is nothing for it to sit still against. Put it back into the master scene if it belongs in the space. If it is meant to be a screen-locked graphic—a caption or tag that always sits on the glass—it must appear identically in both eyes at the reference plane. That's a legitimate deliberate choice, and worth saying out loud so nobody reads the choice as a bug.
Then restore the paired scene and watch the whole short passage, not the friendliest frame. The easy frame is the one with the least motion and the cleanest separation between layers. The informative ones are the cross-over beat where the card passes the product's edge, the frame with the fastest movement, and the cut. After you change anything, re-check the offsets in the frames you touched—a fix that moves the reference plane moves every number in the table.
What the review can honestly carry
There are two complete deliverables here, and they carry different claims.
The stereo review package contains the master scene at a named version, the eye-labeled correct pair, the combined output in the arrangement the stated route expects, the deliberately swapped version clearly marked as a fault, the conditions line, and the viewing route spelled out: which file, on which display, in which mode, with what, in front of whom. Alongside it, a scope statement in plain sentences—what was checked (eye identity, frame sync, offsets against the prediction, occlusion in each eye, one scene, one route, the frames you actually watched) and what was not (other displays, other viewers, comfort, production scale, other scenes, anything about how this behaves at broadcast length).
The flat explanation contains an annotated frame from one eye, a top-down plan showing the camera position and the three layer distances with the reference plane drawn as a vertical line at 120 units, the offset table above, and motion notes marking where the card crosses the product. It carries the intended layering, the occlusion order, the timing, and the composition—enough for a client to judge whether the idea reads. On the artifact itself, not buried in the email, it says: this is a flat explanation of an intended depth arrangement; it does not show the two-eye view.
Compare whichever one you send against the same passage, and be precise about the difference. The flat version can explain the intent of the depth arrangement; it cannot establish the two-eye experience, no matter how carefully it is drawn. The stereo sample, if the route works, shows what the pair does on that route. If reliable stereo review isn't available—no headset, no stereo display, a conference room where the signal collapses to one stream—send the flat explanation, state that the stereo test is unresolved, and don't let the word "stereo" appear in the deck as though something had been demonstrated.
One successful playback is not compatibility. It is one file, on one display, watched by however many people were in the room, which is a real result and a narrow one. Display compatibility and comfortable viewing need the checks a specialist would run, and those are not the checks this method performs. Keep the claim scaled to the evidence: "checked on this display, by this route, on these frames, in this scene" is a sentence you can defend in a review. "Looks great in stereo" is not.
Both packages are finished answers. The one that isn't is a depth-looking still with no pair, no route, and no stated limit on what it proves—and that one is easy to mistake for the confident option precisely because it never has to say what it left out.
Frequently asked questions
When is a two-eye stereo mockup the right deliverable?
When you know what the recipient will open it in and how they will look at it as a pair. If that route is unknown, write a flat breakdown, mark it flat on the artifact, and leave the binocular question open rather than sending a pair nobody can view.
What does a depth-looking flat frame establish?
It can show intended layering, occlusion order, timing, and composition, but it cannot establish the two-eye experience. Shallow focus, wide lenses, vanishing perspective, and staggered layers may read as depth in one image while containing no information about a second view.
Why build one master scene and derive both eyes from it?
To guarantee that both eyes contain the same objects, positions, times, and materials. Hand-editing one eye turns a pair into two similar shots. Label eye identity in filenames or a slate, and do not treat this as reconstructing a second view from a single photograph or render.
How are separation and convergence different?
Separation sets the size of every offset in the scene; doubling it doubles every offset. Convergence sets the zero-offset reference plane. With parallel cameras that plane sits at infinity, so it must be chosen by rotation or horizontal shift, and the two geometries are not interchangeable.
What should the review package include, and what is its limit?
It should include the master scene version, eye-labeled correct pair, the combined output for the stated route, a deliberately swapped fault, a conditions line, and the viewing route spelled out, plus a scope statement of what was and was not checked. One successful playback is not compatibility; display and comfort checks remain separate.