Describe a Treatment's Visual Argument for a Nonvisual Reader
Describe a Treatment’s Visual Argument for a Nonvisual Reader
The easiest description to write is a list of what is in the picture. It is also the least useful thing you can hand a reader who can't see the page.
A station at night. Two figures. A bench. A long coat.
Every item is accurate. None of them says that the two figures stand at opposite edges of the frame with the entire width of the platform between them, that the bench stays empty in all three drawings, or that the third frame pulls back until the distance separating the figures becomes the largest object in the picture. Those are the facts the sequence was built on. A list leaves them out, and adding more nouns does not put them back.
The difference between a description and an inventory is not length or polish. It is whether the text has a subject. An inventory is a set of possessions. A description takes a position on what the image is doing, because the treatment page has already taken one: the image is there to make an argument that the prose alone could not carry. Three moves get you most of the way there, plus a review that tells you whether you got anywhere.
Name the question this image answers
Every image in a treatment exists to answer something. A borrowed still fixes a colour relationship or proves that a scale reads at all. A photograph chosen for texture establishes what a surface will have to feel like. A storyboard panel drawn for the project settles a staging. Captions often name the job — reference, colour only; lighting mood; frame 2 of 4 — but the caption is the label on the argument, not the argument.
So write the argument down before you describe anything. One sentence, in your own working notes rather than in the shipped text:
Frame 3 is here to show that the distance between the two figures is a quantity, not a mood.
The test for that sentence is substitution. Swap in a different image. If the page's claim survives the swap, you have found the decorative content, not the argument, and you can describe it briefly and move on. If the page collapses — if the emptiness of the platform is the claim, and any other station would break it — then you know which details are load-bearing. The empty bench is load-bearing. The bench's design is not, unless someone in the meeting will ask about it.
Some images answer small questions, and their descriptions should be one line. A colour chip does not need a paragraph, and a paragraph attached to a colour chip teaches the reader to skim the descriptions that matter. The proposition test tells you where to stop as well as where to start.
Describe relationships rather than listing possessions
The real work happens in prepositions and quantities. Near, behind, between, above, at the edge, cut off by the frame, entering from the right, occupying a tenth of the width. These are the words that survive translation into somebody else's mind.
Numbers are often false precision — nobody measuring a JPEG can tell you it's 9.4 percent. Legibility is the more honest unit. In frame 1 the figure is small but readable: you can see a coat and the direction the figure faces. By frame 3, both figures are too small to read as people rather than as marks, and the interval between them, with the empty bench near its middle, holds most of the frame's width. That is a scale description that a reader can act on, because it says what the scale denies them.
For a sequence, describe each frame against the one before it, and keep the words for unchanged things unchanged. The bench is empty in frame 1. In frame 2 it is still empty — that word has to come back, so the reader can hear that nothing about the bench moved while something else did. In frame 3 it is a small mark inside the interval, no longer furniture but a unit of measurement. Reach for a synonym for "empty" in that third sentence and you break the chain the reader was using to track distance.
Then hold the line between observation and interpretation. "The wide frame gives its space to the gap rather than to either person" is a reading of the composition, and it is defensible from what is drawn. "The figure feels abandoned" is an interior state no drawing can establish. A description that asserts it has replaced the reader's experience with the writer's, and the reader can no longer disagree.
Absence needs the same care. A panel with no other passengers and no train in it shows what was drawn. It does not show that the platform is empty in the finished film. If the treatment says the shot is to be kept clear of traffic, that is a proposal, and it belongs to a different sentence than the drawing does. The distinction sounds pedantic until a reader builds the scene wrongly: someone who takes the drawn emptiness for a production plan hears isolation as a fact, while someone told "the treatment calls for no other traffic while the figures are this far apart" hears a decision they can argue with.
Duration is the last trap. A still cannot hold. "The camera lingers" and "we stay on the wide" describe time, and time is not in a panel. If the treatment's own text says the wide is held, the description can say so and should credit it to the text. If the treatment doesn't say it, the writer has just added a directing choice. That may even be the right choice, but it is not the writer's to add silently.
Separate identification from fuller explanation
W3C WAI's tutorial on complex images draws the distinction treatments actually need: a brief identification of the image, and then a fuller account of essential content when composition, scale, relationships, or a trend across a series cannot be carried by the short version. The guidance addresses images on web pages. It is not a conformance statement for a tagged PDF or an exported deck, and it promises nothing like an equivalent experience of seeing. Borrow the shape of the distinction; work out the implementation in the format you are actually shipping.
That gives you two registers.
Every image gets an identifying line, including the ones with no argument in them. Frame 2: same view, second figure entering at the right edge. That is orientation, and it is what lets a reader follow the sequence rather than receive a series of disconnected pictures.
Where the argument lives in relations, or in the change from one frame to the next, the fuller account goes with it. One paragraph is usually enough for three frames, and more than that usually means the writer is compensating for a missing relationship somewhere else.
Keep the treatment's own qualifications. Reference only. Mood. Not final. Detail of frame 4. Compare with the opening. Each one changes what the reader should take the image as proving, and a description that drops them has quietly made a mood photograph into a production plan. If the page offers an image as a comparison, describe the comparison rather than describing each half and leaving the reader to guess at the join.
Do not duplicate the body copy. If the prose beside the image already says the scene is about distance, the description's contribution is the visual evidence for it: which frame shows how much distance, and what changed between frames. Restating the claim in fresher adjectives gives the reader nothing to hold and makes the description look longer than the image warranted.
Where the fuller account sits matters as much as what it says. A treatment that runs one long passage under frame 1 to cover the whole sequence leaves a reader who reaches frame 3 with nothing but "image, frame 3" — the explanation is three figures back and the association is broken. Sequence-level accounts belong with the sequence; per-frame lines belong with their frames. In a web delivery the longer text sits adjacent to the image or on a page the image links to. In a PDF the figure needs to carry the text and the reading order needs to reach it. In a slide deck, notes are reachable by some readers and not by others, and a text box laid over the artwork is announced in whatever order the export decides. None of that is settled by well-written prose in a draft document.
Check whether the argument survives without the picture
Review with a nonvisual reader and ask for reconstruction, not opinion. Not is this clear, but: in frame three, where are the two people, and what changed since frame two? A reader who answers "each at an edge, farther apart than they looked, with the empty bench between them" has the argument. A reader who can only say "there's a platform and two figures" has the identification and not the claim.
When the reconstruction comes back flat, the repair is another relationship, not another noun. Adding a bench, a lamp, a shelter, a timetable board thickens the inventory. Adding the bench sits nearer the middle of the interval than to either figure gives the reader something to hold. Fix relations first, then decide whether the object list needs any help at all.
Two failures are worth keeping apart, because they call for opposite responses.
The first is omission. The reader cannot tell whether a train arrives in the last panel, whether these two are the only people on the platform, whether the drawings are final. Where the treatment text settles those questions, the description carries the answers. Where it doesn't, the gap belongs to the treatment, not the description — and a writer who fills it has manufactured certainty on the director's behalf.
The second is deliberate withholding, and it is not the same as an oversight. These three panels stop before anyone moves. That is a choice, and a reader who knows the withholding is intentional can participate in it. A reader who suspects the writer forgot to mention whether the figures meet will ask for the missing information and lose the effect on the way. The way to keep the two apart is to say what the panels do and do not establish: the three drawings show no movement between the figures; whether they close the distance is not answered here. That sentence is both accurate and an acknowledgment that the film's answer does not live on this page.
One caution about the review itself. A description read back in plain text is not evidence that the delivered file exposes it. Whether the PDF's figure carries alternative text, whether the reading order reaches it, whether a screen reader announces the artwork instead of the description — those are questions about the file that ships and the technology used to read it, and they can only be answered in that pairing.
What the description can and cannot do
Used well, a description gives another route to the meaning: not the experience of seeing, but the argument the seeing was arranged to make. What follows is the three-frame sequence again, with the drawn material and the proposed material kept on separate lines.
Frames 1–3 (drawn). Frame 1: a night platform; one figure stands at the far left, the coat cut by the frame edge; the platform runs away to the right past an empty bench with nothing else on it. Frame 2: the same view; a second figure enters at the far right edge; the bench stays empty and nothing else in the frame changes. Frame 3: the same platform pulled wide, until each figure is a small mark near an edge and the space between them — the empty bench somewhere near its middle — fills most of the width. Neither figure is large enough in frame 3 to read as a person in more than outline.
Proposed in the treatment text, not visible in the drawings. The wide in frame 3 is to be held. The camera does not move during the hold. No other passengers and no train enter the shot while the two figures are that far apart. The stopping short of contact is deliberate.
Not decided by the drawings. What either figure wants, and whether the crossing happens.
A reader working from that account can follow the sequence, can tell what the panels establish from what the text proposes, and can put the director's own question back to the page: does the film need the distance, or does it need the hold on the distance? An inventory of a platform, two figures, and a bench gets nobody to a question worth asking. It is accurate, and it finishes a page before the argument starts.
Frequently asked questions
Why isn't a list of objects enough?
A list can be accurate and still omit everything the sequence was built on. 'A station at night. Two figures. A bench. A long coat' does not say that the two figures stand at opposite edges with the entire width of the platform between them, that the bench stays empty in all three drawings, or that the third frame pulls back until the distance separating the figures becomes the largest object in the picture. Those are the load-bearing facts. Adding more nouns does not put them back. A description takes a position on what the image is doing; an inventory is only a set of possessions.
How do I know which details are load-bearing?
Write the argument down in one sentence in your working notes, then run a substitution test: swap in a different image. If the page's claim survives the swap, you have found decorative content; describe it briefly and move on. If the page collapses—if the emptiness of the platform is the claim, and any other station would break it—then you know which details are load-bearing. The empty bench is load-bearing. The bench's design is not, unless someone in the meeting will ask about it. Some images answer small questions, and their descriptions should be one line; a paragraph attached to a colour chip teaches the reader to skim the descriptions that matter.
How should I handle a sequence of frames?
Describe each frame against the one before it, and keep the words for unchanged things unchanged. The bench is empty in frame 1. In frame 2 it is still empty—that word has to come back, so the reader can hear that nothing about the bench moved while something else did. In frame 3 it is a small mark inside the interval, no longer furniture but a unit of measurement. Reaching for a synonym for 'empty' in that third sentence breaks the chain the reader was using to track distance. Work in prepositions and quantities: near, behind, between, above, at the edge, cut off by the frame, entering from the right, occupying a tenth of the width. Use legibility rather than false numeric precision, and keep observation separate from interpretation: 'the wide frame gives its space to the gap' is a defensible reading of the composition, while 'the figure feels abandoned' is an interior state no drawing can establish.
What is the difference between identification and a fuller explanation?
W3C WAI's tutorial on complex images draws the distinction treatments need: a brief identification of the image, then a fuller account of essential content when composition, scale, relationships, or a trend across a series cannot be carried by the short version. Every image gets an identifying line, including the ones with no argument in them, so the reader can follow the sequence rather than receive disconnected pictures. Where the argument lives in relations or in the change from one frame to the next, the fuller account goes with it. Keep the treatment's own qualifications—reference only, mood, not final, detail of frame 4, compare with the opening—because dropping them can quietly make a mood photograph into a production plan. Do not duplicate the body copy; the description's contribution is the visual evidence. Placement matters: sequence-level accounts belong with the sequence, per-frame lines with their frames; in a PDF the figure needs to carry the text and the reading order needs to reach it, and in a slide deck notes are reachable by some readers and not by others. The guidance addresses images on web pages; it is not a conformance statement for a tagged PDF or an exported deck.
How do I review and fix a description?
Review with a nonvisual reader and ask for reconstruction, not opinion. Not 'is this clear,' but: 'In frame three, where are the two people, and what changed since frame two?' A reader who answers that each is at an edge, farther apart than they looked, with the empty bench between them has the argument. A reader who can only say there is a platform and two figures has the identification and not the claim. When the reconstruction comes back flat, the repair is another relationship, not another noun: adding a bench, a lamp, a shelter, and a timetable board thickens the inventory; adding that the bench sits nearer the middle of the interval than to either figure gives the reader something to hold. Keep omission and deliberate withholding apart. Where the treatment text settles whether a train arrives, whether these are the only people, or whether the drawings are final, the description carries the answers; where it does not, the gap belongs to the treatment, not the description. For withholding, say what the panels do and do not establish. Also, a description read back in plain text is not evidence that the delivered file exposes it; whether a PDF's figure carries alternative text, whether the reading order reaches it, and whether a screen reader announces the artwork instead of the description can only be answered in that pairing of file and technology.