Skip to content

Pitch a Forced-Perspective Reveal the Reader Can Actually Understand

Advertising

Pitch a Forced-Perspective Reveal the Reader Can Actually Understand

A forced-perspective beat in a commercial treatment is two claims, not one. The first claim is what the viewer will believe about the relationship between things—that the mug dwarfs the person walking toward it, that the two people are standing together, that the car is parked on the mountain. The second claim is what changes their mind.

Write only the second claim and you have written a punchline with no setup. Whoever reads the treatment—client, producer, director—will either guess the first half wrong or quietly conclude you don't know what the first frame is doing. The whole job of the treatment line is to make both readings visible in language, before anyone builds anything.

Everything below about the mug is my own arrangement, worked out on paper. It has not been shot, tested, prevised or shown to a camera department. Treat it as a proposal you can hand to someone who can test it.

State the false reading as a sentence a stranger could redraw

The first draft of most forced-perspective treatments reads like this: "We open on a surprising forced-perspective shot of the product." That sentence tells the reader a technique was used. It doesn't tell them what the audience will think they're looking at, which is the only thing the rest of the spot has to work against.

So write the inference as a belief about a relationship:

A mug stands on the street at the height of the rooftops. Far down the road, someone walks toward it.

A reader who has never seen a storyboard can now redraw the composition. That's the test. If your sentence can't survive being drawn by someone who only read it, it isn't yet a setup—it's a mood.

The flatness is deliberate. The sentence isn't a claim about the mug; it reports what the frame will make someone believe, in the terms they will believe it. What converts that report into a belief is the correction, which the treatment therefore owes its reader as plainly as it owes them the setup.

Two habits make this easier.

Keep the belief in the viewer's grammar. The mug isn't towering; from where the camera sits, it towers. So the model line stays flat—"a mug on the street at roof height" is a report of the reading, not a property of the mug. Hedging is available and sometimes right: "from this angle the mug appears to stand at roof height" tells a nervous reader up front that the size belongs to the view. It also tells them the frame is arranging something. Spend it deliberately rather than by reflex.

Never let the illusion contaminate the product claim. Write "the mug reads as though it were two storeys tall," not "a two-storey mug." The first is a description of the frame. The second is a claim about the product, and it's false—which matters later, when somebody in the room starts talking about how impressive the capacity must be.

Where things actually are, and what the frame has to hide

Apparent size on screen is close to a simple ratio: the object's real size divided by its distance from the camera. That's the whole engine. You don't need to derive anything technical to use it, but you do need the two distances in your head before you can say what the first frame shows.

Here's an illustrative arrangement, chosen for round numbers:

  • Mug: 0.09 m tall, standing on a windowsill.
  • Camera: about 0.35 m from the mug, at roughly sill height.
  • Friend: 1.75 m tall, walking along the pavement about 25 m away.

The mug's angular height works out to roughly 0.26 radians; the friend's to about 0.07. The mug therefore covers about three and a half times the friend's height in frame. If a viewer reads those two figures as sitting at the same distance, the mug is being described as something around six metres tall—about two storeys.

That is the entire illusion. A nine-centimetre object and a person twenty-five metres apart, compared as though they shared a plane.

what the lens sees (proposed)

  +-------------------------------------------+
  |                                           |
  |   [ MUG ]                                 |
  |   fills the left third,                   |
  |   foot meeting the far kerb                |
  |                            . (friend)     |
  |                            small, right,  |
  |                            walking in      |
  +-------------------------------------------+

Now the part that's actually design work: the frame must exclude every object that sits at the mug's own distance. The windowsill's front edge. The window frame. The wall below it. Any café furniture, bin, bollard or kerbstone close enough to be measured against the mug. The reason the mistake is available at all is that nothing in frame shares the mug's depth, so the viewer has no honest comparison to make and borrows the only one on offer—the person far away.

That's why the excluded list is as important as the included one, and why a treatment that only describes the mug is missing the actual instruction.

Diagram it, even crudely. A plan view and a frame view, side by side, will surface an argument faster than a paragraph:

plan view (proposed, not to scale)

  camera ── 0.35 m ──► [mug on sill]     sill, wall, frame: OUT OF FRAME
    │
    └────────────── ~25 m ──────────────►  friend, walking toward camera

Then mark the alignment as a question rather than a decision. The one I'd flag first: can the mug's foot be made to meet the far streetline while the sill stays below or behind the frame edge? That alignment is what sells the mug as standing on the road at roof height, and it's also the thing most likely to fight the requirement to hide the sill. I don't know the answer. Neither does a drawing.

A second question, just as blunt: can a subject a third of a metre from the lens and a subject twenty-five metres away both hold acceptable focus in the same frame? That's a camera-department question, not a treatment question, and the honest move is to write it into the treatment as an unresolved dependency rather than assume it away.

And a third thing that no diagram can settle: whether a person can actually walk into that arrangement and take the mug without the geometry collapsing on them. A persuasive previs frame or a generated still is a picture of an intention. It is not evidence that a body can move through the space and arrive where the story needs them.

Choose the movement that corrects the relationship

The correction is a change of view. The question the treatment has to answer is which kind, because the two kinds behave differently.

A rotation—pan or tilt—moves the frame without moving the camera. Everything in the shot shifts across the screen by the same angular amount, near and far alike. Nothing re-sorts. The illusion survives the move and new content simply enters from the side. If your reveal is "the world widens and you finally see what the mug is sitting on," a pan is the safer engine.

A translation—dolly, truck, crane, a handheld drift—moves the camera's position. Near objects sweep across the frame faster than distant ones. That relative shift is parallax, and it is the strongest depth cue an audience has. It will expose your arrangement earlier than you want—unless you want it. If the beat is "the mug drops back into the plane of the windowsill, and the gap closes," you're spending the illusion to buy the correction, and that's a legitimate trade. Just decide which one you're making. Writers often describe a general "camera move" and leave the whole reveal to a verb that could mean either.

Now name the new information. Not "the camera reveals the truth"—the specific fact that breaks the earlier belief. Here it's the arrival of a same-plane reference. A mug beside a hand at the same distance is comparable; a person beside a person is comparable; a mug beside a walker twenty-five metres away is not. When the sill, the wall and the arriving friend all sit at roughly the lens's own distance, the frame contains an honest comparison for the first time, and the mug shrinks to what it is.

Which gives you the transition's real content: a reference moves into the mug's depth plane. In the version I'd write, the friend has been walking toward camera the whole time, and the pan left finds them already at the sill, their hand entering frame at the same distance as the mug. The hand is the falsifying fact.

Two things to resist.

Don't cut from the impossible frame to the normal one and let the gap stand in for a decision. A cut substitutes the true view for the false one. It tells the audience they were fooled without telling them how, and the correction stops being a beat and becomes a repair. Sometimes that's fine—if the piece only needs the viewer to re-read the first image, a cut does it. But if the correction is the point, the transition is the scene, and it has to be written.

Check the walking distance against your running time. Twenty-five metres at a normal pace is around eighteen seconds. Shorten the distance to twelve metres and the walk is under nine—but the friend is now twice as large in frame, the ratio falls to under two, and the mug reads as three metres rather than six. Distance is the exchange rate between how big the illusion looks and how long the arrival takes. The treatment should show the writer knows there's a trade there, because the client will feel it in the edit.

What the correction buys the commercial

Ask what the corrected relationship lets the viewer understand, and answer it in terms of the situation rather than the trick.

In this setup: the first frame makes waiting look monumental. Something public, architectural, announced to the street. The reveal makes it domestic—a cup left on a sill for one person. The gap between those two readings is the idea. How large an act of care feels from a distance; how small it looks up close.

Compare it with the plain version. A friend walks up to a windowsill, picks up a mug, drinks. That is already a warm, small beat, and it already carries a welcome. So the trick has to add the one thing the plain version can't: the disproportion. If it doesn't, you've got a puzzle attached to a gesture, and the gesture was doing fine on its own.

Two cautions, both commercial rather than visual.

An illusion of size is not a claim about capacity. If this is a product spot, decide what the viewer should believe about the mug when the frame ends—and for this idea, the answer should be nothing about size at all. They should remember that someone thought of them. Any size impression the audience carries away from the first frame is a distortion you manufactured, not a feature.

And be willing to delete the trick. The test is simple: after the correction, does the viewer reinterpret the first frame? If the answer is "they just learn they were wrong," you've written a gag, and the client paid for a gag. If the answer is "the ordinary thing turns out to be the monumental one," keep it, and say so in the treatment so nobody has to guess why it's there.

The version that goes in the treatment

Before. A mug stands on the street at roof height. Far down the road, someone walks toward it. It has been waiting for them.

Reveal. The camera pans left. The mug is on a windowsill at arm's length, and the walker is already there. A hand closes around it and lifts it away.

Meaning. What looked monumental was a cup someone left out for a friend.

What remains genuinely unknown is the geometry. Whether the mug's foot can meet the far streetline with the sill out of frame. Whether a subject at a third of a metre and a subject at twenty-five metres can both hold focus. Whether a pan brings in the sill without a translation sneaking in and giving the depth away early. Whether the friend can cover the distance inside the shot's length. And whether an audience reads the opening frame as "giant mug" rather than "distant mug."

None of that is settled by describing it well. It's settled by someone with a camera and a tape measure, and until they've been asked, the treatment should promise the idea and not the execution.

Frequently asked questions

What two claims does a forced-perspective treatment have to make visible?

One is the false reading: what the viewer will believe about how things relate, such as a mug appearing to dwarf a person. The other is the correction: what changes that belief. Writing only the correction leaves the reader to guess the setup.

How can you tell whether the false-reading sentence is doing its job?

It should be flat enough that a stranger could redraw the composition from the sentence alone. If it only says a technique was used or describes a mood, it is not yet a setup.

Why does the frame need to exclude objects at the mug's distance?

The illusion depends on the viewer having no honest same-depth comparison. If a sill edge, window frame, wall or nearby street object is visible, the viewer can measure the mug against it. Excluding those things leaves the distant person as the only comparison available.

What is the difference between a rotation and a translation in the reveal?

A rotation moves the frame without moving the camera, so near and far shift by the same angular amount and the illusion can survive. A translation moves the camera, so near objects sweep faster than distant ones; that parallax exposes depth earlier. The treatment should choose which kind of correction it is making.

What trade does walking distance create?

Distance sets both the size of the apparent illusion and how long the arrival takes. In the example, 25 metres at a normal pace is about 18 seconds; shortening it to 12 metres makes the walk under 9 seconds, but the friend is twice as large in frame, the ratio drops below two, and the mug reads as about three metres rather than six.

More in Advertising Browse all articles