Skip to content

Animate Depth in a Still Photograph for a Treatment Sample

Advertising

Animate Depth in a Still Photograph for a Treatment Sample

A depth move built from one photograph is a constructed scene. Nothing in it photographs anything. It separates pixels the camera already recorded, slides them past each other, and asks you to supply whatever slides into view. The method is short: pick a photograph with a usable front-to-back split, cut a few planes, keep the camera travel modest, and inspect the areas that the travel exposes.

The uncomfortable part is the arithmetic. One quantity governs both how much depth the move appears to have and how much of the frame you have to invent. The relative shift between two planes is the hole behind the near one. There is no version of this where you collect the depth and skip the hole. Everything below is about spending that quantity deliberately.

This covers plane separation, a camera move, the exposure it creates, and the comparison that tells you whether to bother. It does not cover rebuilding a scene in three dimensions, which is a different project with a different budget and a different set of problems.

Choose a photograph with a useful separation

A usable source has three properties, and they are practical rather than aesthetic.

The first is a cuttable edge. Something in the frame has a boundary you can isolate without spending an afternoon on it: a cup against a wall, a chair back against a window, a hand against a sleeve. Hair, foliage, smoke, and glass are all technically possible and all slow. The move will put that edge under scrutiny for its whole duration, so an edge that survives a close look at one frame has to survive every frame.

The second is a real depth relationship the treatment needs to show. A depth move demonstrates that the space has layers — that there is a near thing, a far thing, and a viewer positioned among them. If the treatment's point is the subject's expression, or a color, or a product's finish, a depth move introduces a second, competing claim about space and asks the viewer to care about it. Add depth because the pitch needs the room to read as a room, not because a still frame felt insufficiently alive.

The third is pixel headroom, and this is the one people skip. A photograph that fills a 1920-wide frame from a 4000-pixel-wide source is being displayed at roughly half size. That surplus is the budget for pushing planes away from the camera, because a plane that sits farther back has to be scaled up to keep its starting size on screen. A photograph that already fills the frame at 1:1 has no such budget, and any plane you place behind the others will be enlarged the moment you start. You will not notice at the first frame. You will notice at the last one.

Keep the untouched file. It is the reference you check every frame against, and on a construction like this the original is the only thing that tells you which pixels came from the scene.

One boundary note, briefly: if the photograph is not yours, the permission question comes before the craft question, and it covers the derived sample as well as the source.

Separate planes and account for what was hidden

A photograph has one surface. There is no second view stored behind anything in it. When you lift a foreground object onto its own plane, you are not revealing the wall behind it — you are revealing that the wall behind it was never recorded. The pixels that appear in that gap are yours, not the scene's. Hold onto that distinction, because it is the whole ethical content of the technique.

Cut only the planes the move needs. Two or three covers most treatment samples: a full-frame background, a middle card, and a near card. There is a useful economy available here — planes that end up at the same depth never move relative to each other, so anything that shares a depth can live on one layer and cost you nothing extra in exposure.

Then decide what fills the gaps, and label the decision. There are three honest options. You can extend adjacent texture by cloning or painting from the surrounding area, which works when the region behind the object is continuous. You can lay in a soft neutral patch that reads as out-of-focus surface, which is a visible admission rather than a claim. Or, for internal review, you can leave the gap marked and unfilled so the team can see exactly where the construction starts and stops. Any of these is fine. The failure mode is doing the first one and then describing the result as detail recovered from the photograph.

Keep the record. Retain an unfilled version of the composition in the project. Make a still with the filled regions outlined. The reason is not paperwork: three weeks later, when someone wants to extend the move by a second, the marked-up version is what tells you which areas cannot be extended.

Make a small camera move through the constructed layers

Two documented behaviors do most of the work. Adobe's help page on 3D layers in After Effects describes layers occupying distinct depth positions that are rendered through the active camera. The page on cameras, lights, and points of interest describes camera settings and movement changing the view of 3D layers, with the active camera determining the output view. Those are the tool's rules, taken from public documentation. Nothing below has been run in a composition for this draft — the geometry is reasoning, and the example further down is a plan rather than a result.

The first rule's consequence is a habit worth forming: preview through the camera that will render. A viewport showing a top view, a custom view, or a camera you animated last week is not the output, and it will happily show you a beautiful move that the render will not produce. Set the active camera before you start judging anything.

The second rule is about scale. Apparent size on screen is proportional to a plane's size divided by its distance from the camera. If you place a plane at twice the distance, it needs roughly twice the scale to hold the same starting composition. So each plane's scale is set by its depth, and the farther-back planes are the ones absorbing enlargement. Planes that come toward the camera get scaled down, which costs nothing.

Now the part that matters. Parallax in a setup like this comes from translation — moving the camera sideways through the arrangement. A camera rotation, by itself, produces no relative movement between planes at all: rotating a pinhole camera shifts every point in the scene by the same angle, so nothing slides against anything else. This is also why a pan over a flat photograph looks flat. It is not a limitation you work around. It is the reason the technique needs a translation, and the reason the translation has to be paid for.

Here is the shape of the payment, in numbers simple enough to hold in your head. Measure your camera travel as a fraction of how much of the world the frame covers at the near plane's distance, and call that fraction t. The near plane slides across the frame by t; the far plane slides by t multiplied by the ratio of the two distances. The gap between those two shifts is what the viewer reads as depth, and it is the same gap you have to fill:

relative shift = t × (1 − near distance ÷ far distance)

Travel (t) Depth ratio Relative shift, as a fraction of frame width
6% 3 : 1 4.0%
6% 2 : 1 3.0%
6% 10 : 1 5.4%
20% 3 : 1 13.3%
3% 3 : 1 2.0%

In a 1920-wide frame, the first row is about 77 pixels of movement and about 77 pixels of unrecorded area to account for. The fourth row is about 256 pixels of each. There is no row where the two columns diverge, because they are the same column.

So if the exposed strip is wider than you can fill, the fix is less travel or less depth separation — not a cleverer fill. Both of those reduce the parallax by exactly the amount they reduce the hole.

Two things follow from that, and they are the practically useful ones.

The first is that the arrangement is scale-invariant. If you take the whole setup — camera, planes, travel — and shrink it toward the camera, the rendered frames are identical, as long as nothing is being enlarged past 1:1. This is the tool for the pixel-headroom problem above: rather than pushing the background plane far away and paying for it in stretched pixels, keep the whole arrangement compact and scale the camera travel down to match. Depth here is relative. Nothing in the image says how big the room is.

The second is where the cost lands. The plane you push back is the one that needs the most enlargement. The fills land by position: a lifted card opens its gap onto whatever layer sits directly behind it, which is often the middle card rather than the background, and only a gap that reaches the background is filled there. Enlargement and unknown content can therefore land on different layers. That is a good reason to keep the background near, and to keep its scale at or under 1:1 for the whole move.

If you want the framing to return to your subject after a translation, recenter with a camera rotation. The rotation adds no relative shift, so it will not disturb the parallax you just paid for.

Inspect the path, not just the opening frame

The first frame is the one frame you already know is fine, because it matches the photograph. The informative frames are at the ends. Scrub the whole move and look for:

  • Exposed holes. Where a lifted card has moved far enough to reveal area with no recorded source.
  • Matte edges. A cutout carries a thin halo of whatever it was originally photographed against, mixed into its antialiased boundary. Over a new background, that mixture is wrong, and it reads as a glowing outline at this scale.
  • Overlap changes. If two cards swap which one is in front partway through, the depth order is wrong or the move is larger than the arrangement supports.
  • Plane edges entering frame. The background plane is a finite rectangle. A move that runs off its side does not show you the wall continuing; it shows you the edge of your own layer.
  • Scaling that reads wrong. A card that grows slightly too fast reads as having been closer to the camera than the scene allows. This is subtle and it is also the kind of thing a viewer registers as "off" without naming.

Find the frame with the largest relative shift and look at that one first. Then decide between two repairs. You can patch the specific boundary — targeted, cheap, and appropriate when the gap crosses a texture you can extend. Or you can shorten the travel, which reduces the exposed strip roughly in proportion: halve the travel, halve the strip. On a treatment sample, the second option is usually the better one, because a slightly shallower read is a smaller loss than a fill that invents structure.

And that is the distinction worth carrying out of this section. A 250-pixel gap in a smooth wall is mostly work. A 60-pixel gap that runs across the continuation of a table edge, a shelf line, or a shadow boundary is a different category of problem: you are no longer extending texture, you are inventing structure. Size matters less than what the gap cuts through.

Compare constructed depth with an ordinary pan

Build the comparison before you commit. Take the intact photograph, set the same starting framing and the same duration, and animate an ordinary pan across it to roughly the same subject position. Then watch both.

The pan's pixels are all recorded. Every frame of it is a window onto something the camera actually saw, and its implied claim stops at the edge of the photograph. The depth version's frames contain pixels that were never recorded. Where the fills are seamless, a viewer cannot tell which ones; where the fill is one of the visible admissions — a neutral patch, or a marked gap — the viewer can see the construction for what it is. That is the real difference between the two — not the motion, not the polish. The depth move is making a claim about space that the source cannot support, and only a seamless fill keeps that claim out of sight.

So the useful question is: what does each version ask the viewer to believe? If the treatment's point is that the space has layers — a table in front of a person in front of a wall — the constructed depth earns its cost. If the point is the object itself, the pan carries it, and the pan is cheaper in a way that will matter later: it is one animated parameter, whereas the depth version has planes, scales, a camera path, and fills. Change the travel by a few percent and you revisit every fill. Change one number on the pan and you are done.

Drop the depth move if it changes what a viewer thinks exists — if the exposed area resolves into a doorway, a person, a window, or any object the photograph does not contain. That is no longer an animation decision. It is a claim about the scene, and it belongs to the production, not to a sample.

The example this piece plans but has not built

No photograph was made for this draft, no composition was built, and no frames were rendered. What follows is the setup, with the numbers you would check before spending time on it. When it is built, everything marked as expected needs to be replaced with what the frames actually show.

The source. One original tabletop still life, shot straight on to the table edge on a tripod, 4000 pixels wide. Three things separated in depth: a cup in front, a bowl in the middle, a textured wall panel behind. The panel matters — it is continuous enough that extending it is plausible, and it gives the gaps something to be checked against.

The composition. 1920 × 1080, a 50 mm perspective camera, three layers. The background plane carries a cleaned plate: the photograph with the cup and bowl removed and the wall panel extended across the areas they covered. The untouched photograph stays in the project as the reference file, not loaded as a layer. The plate is placed at twice the near card's distance and scaled to hold the frame at the start, inside the headroom the source allows. The bowl is cut on a matte at the middle distance. The cup is cut on a matte and placed at the near distance.

The target move. Travel of 6% of the frame width at the near distance. By the relation above, that is a relative shift of 3% and a gap of roughly 58 pixels in the 1920-wide frame, opened behind the cup's trailing edge.

The overextended version. Identical planes, travel raised to 20%. Same construction, same fills, roughly 192 pixels of exposed background — about 3.3 times the area to account for, in the same place, with the same information available to fill it.

The restrained revision. Travel cut to 3%, holding the two-to-one depth. About 29 pixels of gap and about half the parallax in the target version. The move reads flatter. That is the trade, stated plainly enough to argue with.

The ordinary pan. The intact photograph, panned across to roughly the same subject position over the same duration.

The frames to inspect first are the last ones, where the cup has moved furthest against the wall panel, and specifically the moment when the revealed strip reaches the panel's shadow line — the point where extending texture stops being possible and inventing structure starts.

What the sample says about itself

Finish either way. If the depth move survives inspection, export it with a construction note attached — a sentence in the treatment that says the move is built from a single photograph, that the space behind the near object is illustrative, and that any areas filled were filled. If it does not survive, keep the pan, or keep the still.

The test for which one you owe the room is simple enough. Read the sentence aloud. "The depth move is a constructed parallax from a single photograph; the areas behind the near object are illustrative, not recorded." If that sentence would weaken the treatment's argument, then the treatment is leaning on a claim the sample cannot support, and the pan is the more honest piece of work.

Frequently asked questions

What makes a photograph usable for a depth move?

Three practical properties. First, something in the frame has a boundary you can isolate without spending an afternoon on it — hair, foliage, smoke and glass are all technically possible and all slow, and the edge will be under scrutiny for the whole move. Second, the depth relationship has to be one the treatment needs to show, since a depth move introduces a second, competing claim about space. Third, the source needs pixel headroom: a photograph displayed at roughly half its native size has surplus to spend on pushing planes away, and a source already at 1:1 has none. On basis: the geometry described is reasoning and the worked setup is a plan rather than a result — no photograph was made, no composition built and no frames rendered, so its expected numbers need replacing with what the frames actually show.

What single quantity governs both the perceived depth and the work of filling it?

The relative shift between planes — it is the hole behind the near plane, and there is no version where you collect the depth and skip the hole. For camera travel measured as a fraction of how much of the world the frame covers at the near distance, call it t, the relative shift is t × (1 − near distance ÷ far distance). In a 1920-wide frame, 6% travel at a 3:1 ratio is about 77 pixels of movement and about 77 pixels of unrecorded area; 20% travel at the same ratio is about 256 pixels of each. If the exposed strip is wider than you can fill, the fix is less travel or less depth separation, not a cleverer fill.

Why can’t a camera pan produce the same parallax?

A rotation on its own produces no relative movement between planes: rotating a pinhole camera shifts every point in the scene by the same angle, so nothing slides against anything else — which is also why a pan over a flat photograph looks flat. Parallax comes from translation, and the translation is what has to be paid for. A rotation is still useful for recentring after a translation, because it adds no relative shift and so will not disturb the parallax just bought.

What are the honest ways to fill the areas a lifted plane exposes?

Extend adjacent texture by cloning or painting from the surrounding area, where the region behind the object is continuous; lay in a soft neutral patch that reads as out-of-focus surface, which is a visible admission rather than a claim; or leave the gap marked and unfilled for internal review, so the team can see where the construction starts and stops. The failure mode is doing the first and then describing the result as detail recovered from the photograph. Keep the record — an unfilled version in the project and a still with the filled regions outlined — because it tells you later which areas cannot be extended.

How do I inspect the move, and when is a gap more than a matter of size?

The first frame is the one frame you already know is fine, because it matches the photograph; the informative frames are at the ends. Scrub the whole move for exposed holes, matte halos where a cutout’s antialiased boundary carries its original background, cards swapping front-to-back order, the background plane’s own edge entering frame, and scaling that reads slightly too fast. Look at the frame with the largest relative shift first, then either patch that boundary or shorten the travel — halving the travel halves the exposed strip. Size matters less than what the gap cuts through: a 250-pixel gap in a smooth wall is mostly work, while a 60-pixel gap running across the continuation of a table edge, a shelf line or a shadow boundary means inventing structure rather than extending texture.

More in Advertising Browse all articles