Stereo and Flat Versions Need Different Images. Which Commercial Are You Pitching?
Stereo and Flat Versions Need Different Images. Which Commercial Are You Pitching?
A stereoscopic render and the flat export made from it differ by exactly one cue. Composition, movement, grade, the product, the light, the music: identical. That is what makes the mistake so easy. The flat version looks like the same commercial, so the treatment treats it as the same commercial.
When the reveal depends on that one cue, it isn't the same commercial. It's a different spatial sentence, and on a bad day it's a sentence you didn't mean to write.
Everything below is treatment-stage reasoning and one proposed example. Nothing here has been rendered, tested, or reviewed on a stereo display, and no brief owner has signed off on anything.
Name what depth communicates in the decisive moment
Find the beat where the viewer's understanding changes. Not the shot that looks best in board form — the beat that does the explaining.
In the proposed board, there are three positions along the camera axis: a near element (a pane of glass, a hand, a panel of light), a graphic ribbon, and the product farther back. The board wants the ribbon to pass through the gap between the near element and the product, and the decisive instant is the moment the ribbon is in clear space — occluded by nothing, occluding nothing.
What does the depth purchase there? Not prettiness. The claim is that the ribbon belongs to the room the product is standing in. It's a thing in the scene, not a graphic laid over the shot. That is a specific meaning, and it's the one worth protecting. If you can't say what the depth is asserting, you don't have a stereo beat. You have a nice render.
Now separate the two mechanisms, because they get conflated constantly.
A flat image already carries depth information. Occlusion is the strongest of those cues: a nearer shape cutting a farther one. Add relative scale, focus falloff, haze, and motion parallax whenever the camera moves, and a flat picture is not depthless.
But notice what occlusion gives you. It gives you order. If the near element cuts the ribbon and the ribbon cuts the product, a flat viewer has the stacking right, top to bottom. What occlusion never gives you is the measurement — how deep into that gap the ribbon actually sits, how much room separates it from the product, whether the gap is a hand's width or a hallway.
A stereoscopic presentation supplies something none of those cues supply. The two eye images differ slightly in the horizontal position of the same objects, and the visual system reads that difference as placement relative to the screen plane. Images rendered with no such difference sit at the plane; a difference in one direction puts them in front, the other behind. That is the cue the flat export drops. Everything else survives.
So the collapse isn't "three planes become two." It's worse and more specific than that. The board's decisive frame has no overlaps in it at all. In stereo that frame still fixes a relation: the viewer reads the ribbon as behind the near element and in front of the product, and reads which of those two it is nearer. That is placement relative to the scene, not an exact position in the room, and it is enough for the beat. In flat, even that relation goes unstated: a shape against a background, floating at a depth nobody can name, with no familiar size to reason from because the ribbon is a synthetic element with no real-world scale. Across the rest of the shot the ordering may hold up fine. At the beat that matters, flat has nothing to say.
That's the test. The flat viewer isn't missing an effect. They're being told less — or, depending on how the later frames stage the overlap, told something adjacent that reads as a logo stamped on the product.
A useful exercise before you write another paragraph of treatment: describe what the viewer understands at that beat without using the word between. If the sentence only works with between in it, you've located a depth dependency, and you know precisely which beat needs attention.
Test a common action before promising two versions
Two versions sounds like twice the work. Sometimes it is. But there are levers that can pull both versions back onto a shared action, and it's worth trying them before you commit to separate filmmaking.
Staging. Move the reveal so it lands on flat-legible cues. Widen the gap between the near element and the product along the camera axis. Arrange the ribbon's path so its passage produces real overlaps — cut by the near element at one moment, crossing in front of the product at another — instead of crossing open space where nothing reads.
Movement. Stop trying to show between in a single frame. Sequence the occlusion: the ribbon is cut by the near element first, then travels on and passes in front of the product. A span that goes from behind one thing to in front of another must have crossed what lies between, and a flat viewer will infer that gap from the order of events. One frame can't say it. Three seconds can.
An explicit reveal. Let the product answer. The ribbon's passage throws a shadow across the product's face, or spills light over its surface. Shadows and light spill are flat-legible statements that two things share a space, and they're robust: nobody misreads a shadow traveling across a product as a logo pasted on a photograph.
Any of these routes gives you one set of assets, one continuity, one round of approvals, and two legible versions.
There is a fourth option, and it should be offered rather than smuggled in: drop the claim. Let the near element itself move aside, or let the ribbon simply arrive in front of the product and stop pretending it lives in the room. That is a creative decision, not a technical one, and it's cheaper than everything else on this list. It also means the spot is no longer about a graphic inhabiting the product's space.
What the shared routes cost is worth saying out loud rather than quietly absorbing. In a common-action version, binocular depth stops carrying the discovery and becomes atmosphere. The stereo output still looks dimensional; it just isn't doing the argument anymore. That's a real expressive compromise, and whoever is paying for the stereo pass is entitled to hear it before they see the boards. If the whole reason for the stereo version is that held instant of the ribbon hanging in the gap, the common route has removed the reason.
Two further dependencies, both of which are production decisions rather than freebies. The flat version may need the camera to move laterally, because motion parallax is the monocular cue that most directly substitutes for disparity — and a lateral move changes the shot. That substitution has its own price: motion parallax needs angular change over time, so the flat beat often has to run longer, or the camera has to travel farther, than the stereo beat ever needed to. A spot has a fixed length, and those seconds come from somewhere else.
And keep the honest limit in view. Writing out a flat description of the reveal is an excellent way to expose the dependency. It is not evidence about the stereo result. A paragraph cannot tell you how the depth feels, whether the beat reads, or whether the placement holds in a moving shot.
Compare separate passages and their dependencies
If the shared route loses the governing idea, stop trying to share one passage. Give each version its own.
The stereo passage can be the held moment. The camera rests, the ribbon occupies the gap, the viewer's own visual system supplies the assertion that the graphic is in the room, and the beat can be brief because the reading is instant. Nothing has to move for the meaning to land.
The flat passage has to earn the same claim with different machinery. A plausible design: the camera makes a slow lateral move so the ribbon's position shifts against the near element and against the product at different rates; the ribbon's edge is cut by the near element, then clears it and crosses in front of the product; and a soft shadow from the ribbon travels across the product's surface as it goes. That version costs more screen time and a different shot. It buys a reading that survives on a phone.
The shared inventory is short and worth naming out loud, because it's what keeps the two outputs recognizably the same campaign: the ribbon, the product, the near element, the order of the reveal, and the claim the whole spot is making. The unshared inventory is where the two commercials part: the camera behavior at the decisive beat, the length of the beat, the prominence of the shadow or light-spill element, and — if this is a live shoot rather than a render — the depth placement itself, which is settled when the scene is staged or built, not repaired later on a timeline. That placement is a staging question that belongs in the same conversation as the lighting, not a delivery setting someone finds at the end.
Two passages create a question the treatment has to answer, and it isn't the director's to answer alone: which version's experience governs when the two conflict? Ask the brief owner directly. Three criteria usually settle it. Which version will more people actually see — the flat cut travels furthest, onto phones, laptops, seat-back screens, and every surface that isn't a stereo display. Which version the campaign's central claim depends on. And which version you can genuinely test before committing money.
If the answer is "flat governs, and stereo is the premium experience," that's a legitimate and common answer — but it changes the order of work. The reveal gets designed flat first and translated for stereo, which means the held moment in the gap was never the primary idea. If the answer is "stereo governs," then the flat version is a translation and should be built as one, with its own beat rather than a re-export.
It's worth separating this from a neighboring problem so the treatment doesn't blur them. Framing for a headset, where the question is which part of a scene a viewer will turn toward, concerns where attention goes. This concerns what a single unchanged image means. Different question, different fix.
Choose the evidence object each version needs
Two claims are in play, and they need two different objects. Confusing them is how good boards turn into bad approvals.
For intent, use a flat diagram, and make it do the arguing. A top-down plan showing the near element, the ribbon's path, the product, and the screen plane — plus a storyboard strip marking the occlusion order frame by frame — explains the relationship well enough for a producer or a brand lead to say "that doesn't read" before anything is built. Label it as what it is: an explanatory proposal, not a depiction of the result.
For the stereo question, the only useful object is a prepared sample viewed through the actual display route, by someone qualified to judge it, with the routing configured. That last part is not pedantry. Epic Games' Unreal Engine 5.8 documentation for stereoscopic rendering with nDisplay — opening description and rendering-format sections, notes dated September 18, 2026 — describes a route that generates separate eye images and states that the display hardware has to interpret and route those images appropriately (https://dev.epicgames.com/documentation/unreal-engine/stereoscopic-rendering-with-ndisplay-in-unreal-engine?lang=en-US). That is a statement about how images are produced and delivered. It says nothing about whether your ribbon reads, whether the depth sits right, or whether anyone will enjoy watching it.
Record the conditions whenever a stereo sample exists: which display, which route, who viewed it, what they were asked, and what they reported. A sample viewed on one system is evidence about that system and nothing wider.
And keep a short list of things not to present as proof. An anaglyph is a demonstration format with artifacts of its own, not the delivery route, so it cannot stand in for one: it says nothing about correct playback, comfort, broad compatibility, or whether the production can hit its date. A side-by-side pair deserves a narrower objection, because it is a real delivery arrangement on many displays rather than a demonstration-only format. The objection is to the unrouted pair on an arbitrary screen: that does not demonstrate the configured route, correct playback, or comfort. It becomes evidence only when the designated display routes it as intended, and a qualified viewer records the conditions. A single eye's frame is a perfectly good composition reference and a flat picture of a stereo design; it cannot show the between. None of these objects substitutes for the test you haven't run.
What to say in the room
Name three things plainly and the pitch holds together.
The governing version: which experience the campaign's claim depends on, and which one more people will see. The shared proposition: the ribbon and the product occupy one continuous space, and that is why the graphic isn't a caption. The deliberate difference: the stereo version holds the gap, the flat version crosses it in movement, and the two beats are built rather than exported.
Then hand over one question for the stereo review, narrow enough to have an answer. Not "does the depth look good." Something like: at the reveal beat, does a viewer read the ribbon as occupying the space between the near element and the product, or as a graphic laid over the product? That question is the commercial. Everything else is delivery.
Frequently asked questions
What cue does a flat export drop from a stereoscopic render, and what does that cue supply?
A stereoscopic presentation supplies a slight horizontal difference between the two eye images, which the visual system reads as placement relative to the screen plane. Images with no such difference sit at the plane; difference in one direction puts them in front and the other behind. That cue is what the flat export drops.
Why is the collapse in the decisive frame not simply three planes becoming two?
The board's decisive frame has no overlaps. In stereo it still fixes a relation: the ribbon reads as behind the near element and in front of the product, and nearer one of them. In flat, even that relation goes unstated, leaving a shape against a background at a depth nobody can name.
What shared-action levers can make both stereo and flat versions legible?
Staging can move the reveal onto flat-legible cues, such as widening the gap or arranging real overlaps. Movement can sequence occlusion so the ribbon is cut by the near element and later crosses in front of the product. An explicit reveal can use a shadow or light spill across the product. The cost is that binocular depth becomes atmosphere rather than carrying the discovery.
What does a lateral move cost if the flat version uses motion parallax?
Motion parallax is the monocular cue that most directly substitutes for disparity, but a lateral move changes the shot. It needs angular change over time, so the flat beat often has to run longer or the camera has to travel farther, and those seconds come from somewhere else in a fixed-length spot.
What evidence objects fit the intent question and the stereo question?
For intent, use a flat diagram or storyboard strip labeled as an explanatory proposal, not a depiction of the result. For the stereo question, the useful object is a prepared sample viewed through the actual display route by a qualified person with routing configured. An anaglyph cannot stand in for the delivery route, an unrouted side-by-side pair does not demonstrate the configured route, and a single eye's frame cannot show the between.