Skip to content

Make a Descriptive Transcript of a Commercial Pitch Sample

Advertising

Make a Descriptive Transcript of a Commercial Pitch Sample

A pitch sample usually arrives with two text artifacts attached: a script, and sometimes a caption file. Neither is a descriptive transcript, and the gap between them widens as the clip gets shorter. A 38-second clip with one spoken line produces a caption file of three or four cues. Those cues are correct — timed properly, on screen when they should be — and they still leave a reader blind to nearly everything the clip is doing.

A descriptive transcript is the version you read straight through instead of watching. It carries the spoken words, the sounds that carry meaning, and the visual facts the sound does not supply, in the order the film reveals them. For a pitch sample, that difference usually decides whether the reader understands the argument at all.

The example below is invented, not a record of anything. I have fixed its details deliberately, because the comparison only teaches something if the facts hold still. The method is the one you would run on a real file.

The sample, and two versions of its text

The fictional clip is handoff-v4, 38 seconds, a pitch sample for a courier handoff. An office lobby. A glass door to the street on the left, a reception desk on the right, two elevators at the back. One parcel, one spoken line.

What happens, in order: street noise through the glass; a door chime as a courier in a green jacket enters carrying the parcel; he crosses the lobby and speaks to the receptionist; she does not answer; he sets the parcel on the counter with the label facing away from the camera and looks down at his phone; a man in a grey coat walks over, picks up the parcel and carries it to the elevator; the label comes into view as he turns; the doors close; a close-up shows the courier's phone reading Delivered; the courier leaves, and the chime sounds again.

Here is what a caption file for that clip contains. This is the whole thing.

[door chime]
I'll just leave it with you, then.
[door chime]

And here is a descriptive transcript of the same clip.

handoff-v4 — descriptive transcript
38 seconds, no narration, one spoken line.
Description in plain text. Speech in quotation marks, with the speaker named.
Text visible on screen or on objects in [brackets and caps].

0:00  Lobby of an office building, daytime. Glass door to the street on the
      left. Reception desk on the right, staffed. Two elevators at the back.
      Street traffic audible through the glass. The receptionist is typing
      and does not look up.

0:05  A door chime. A courier in a green jacket and cap enters through the
      glass door carrying one parcel. He crosses the lobby toward the desk.
      A man in a grey coat, wearing a visitor lanyard, stands near the
      elevators with his back to the room.

0:14  Courier, at the desk: "I'll just leave it with you, then."
      The receptionist does not answer.

0:18  The courier sets the parcel on the counter with the label facing away
      from the camera, then looks down at his phone.

0:21  The man in the grey coat walks to the counter, picks up the parcel and
      carries it toward the elevators. The receptionist continues typing.
      The courier's head stays down, facing the phone.

0:23  The parcel's top label becomes legible as the man turns:
      [HOLD — VERIFY ID]

0:26  The man steps into the elevator. The doors close.

0:30  Close-up, the courier's phone: a green check mark and [DELIVERED].

0:34  The courier walks back out through the glass door. The door chime
      sounds again.

Same words. Nothing spoken has been added — the transcript contains exactly the one sentence the clip contains. What changed is that a reader now knows who says it, that nobody answers, that the parcel does not stay on the counter, that a third person takes it, and that the label said HOLD — VERIFY ID. That is the pitch. The caption file has the words of a delivery being left with reception; the transcript has a delivery that went somewhere else while the person responsible was looking at his phone.

Notice also what the transcript refuses to do. It does not say the courier fails to notice. It says his head stays down, facing the phone. A reader can conclude the rest. Description that reports what the camera shows keeps the transcript checkable; description that reports what a character is thinking is an interpretation wearing a transcript's clothes.

Fix the sample and the transcript's job first

Pitch clips multiply. handoff-v4, handoff-v4-trim, handoff-v4-client, handoff-final-2. Before writing anything, put the exact filename, duration and date in the transcript's header, the way the example does with handoff-v4 and 38 seconds. A transcript attached to the wrong cut is worse than no transcript, because a reader will trust it.

Then decide who reads it, because that changes how much you write. A client who cannot play the file needs the visual detail. A colleague taking notes on a call mainly needs the speech and the beats. A reviewer checking whether the pitch lands needs the reveal order intact. One transcript can serve several of them, but only if you know which one you are writing for.

Then take inventory. For handoff-v4 the inventory is small and worth writing down anyway:

  • Spoken words: one line, at 0:14.
  • Meaningful non-speech: two door chimes, one at each threshold.
  • Text on objects or screens: the parcel label, the phone status.
  • Visual events: entry, the approach to the desk, the set-down, the pickup, the elevator, the exit.

That list is the transcript's job description. Anything not on it is either decoration or a claim about the film's meaning, which belongs in your pitch notes rather than inside the transcript.

W3C's Web Accessibility Initiative draws a usable line here: a basic transcript covers speech and relevant non-speech, while a descriptive transcript adds the visual information that a listener would otherwise miss. That is the distinction to work from. It is guidance about making media alternatives, not a certification, and one transcript is not an access strategy on its own.

Correct the words against the recording

The approved script is not a source for the transcript. Neither is the version of the line in the treatment, and neither is what everyone in the room remembers the courier saying. Pitch drafts get rewritten in the edit, lines get dropped for time, and a sample that reads confidently on paper can arrive in the cut as a half-sentence. The recording is the authority.

If you start from an existing caption file or an automated transcription, treat its output as a draft with errors in it. Automated systems handle proper nouns worst — product names, client names, place names, anything coined for the campaign — and they handle the words that matter most in a pitch clip, the specific claims, with the same casual confidence as the filler. Read the result against the audio with your finger on the timeline.

Speaker identification matters more than it looks. handoff-v4 has one voice, but two people are on screen, and a reader who assumes the line belongs to the receptionist understands the clip backwards. When there is more than one speaker, label each turn and keep the labels consistent all the way through.

When you cannot make out a word, write that. A bracketed [unclear 0:14] is honest and checkable. Filling the gap from the script, or from what the sentence was obviously going to be, produces a confident line that nobody said.

One last caution: silence is not an invitation. The receptionist's non-answer at 0:15 is a filmed event, but it is not a beat of dialogue, and writing [no response] inside the speech line invites a reader to hear an exchange that never happens. Say it in description, where it belongs.

Supply what the audio alone cannot establish

This is the part that makes the transcript descriptive rather than merely accurate.

Add the visual facts that change what the words mean. In handoff-v4 that is three things: who receives the parcel, that the receptionist never engages, and what the label requires. None of them are audible. All of them are the point.

Add onscreen and printed text. The parcel label and the phone status are text a hearing viewer reads; a transcript that omits them has dropped a channel, not a detail. Mark such text so a reader can tell it apart from description and from speech — brackets, caps, a stated convention in the header, whatever you choose, applied consistently.

Do not repeat what the speech already establishes. If a line names the product, you do not need a description sentence naming it again as well.

Preserve the reveal order, and this is where transcribers who have watched the clip five times do the most damage. In handoff-v4 the label is legible at 0:23 and not before, because the parcel sits with its label away from the camera until the man turns. A writer who has read the label a dozen times is tempted to mention it at 0:18, where it would be tidy. That single move turns a delivery into a warning and drains the pitch of its tension. Place information where the film places it, even when you know what is coming.

Keep observable and interpretive language distinct. "The receptionist continues typing" is observable. "The receptionist ignores him" is a judgment. "The courier's head stays down, facing the phone" is observable; "the courier does not notice" is a conclusion the reader should be allowed to reach alone. Where a visual fact needs an interpretive gloss in a pitch context, you can add one — but keep it visibly separate from the transcript proper, or mark it as your note.

Make it readable, not a dumped caption file

The caption file's line breaks exist to fit a screen and a reading speed at a specific moment. Neither of those constraints applies to a transcript, so inheriting them produces a column of broken fragments that reads like a film that stutters.

Group action into short paragraphs built around events. Label speakers consistently. Keep a timestamp where it helps someone find a moment — the threshold chime, the handoff, the label becoming legible — and drop the ones that only mark where a cue used to begin. The transcript above uses an anchor at each visual beat because the clip is 38 seconds; a four-minute pitch sample would do better with anchors at section boundaries and a short contents list at the top.

Length is not the measure of success. The transcript here is roughly ten times the length of the caption file, but a transcript that is long because it describes the lobby wallpaper at 0:00 and short because it skips the label at 0:23 has failed in both directions. The test is simpler: read your draft with the video paused. Does the reader know what kind of moment just happened, and does the next line depend on information they already have?

Deliver it beside the right version, then check it again

The transcript needs to live next to the sample it describes, labelled with the sample's identity, in a form a reader can actually use — a page, a plain text file, a document attached alongside the video. W3C's guidance makes the same point from the other direction: a transcript nobody can find does not do its work.

Then recheck. When the clip is re-cut, the transcript ages badly and silently. A tightened edit can drop the spoken line, move the handoff before it, or change the label from HOLD — VERIFY ID to something shorter, and every one of those changes invalidates a passage. Retest the transcript against the finished cut, not against the cut you transcribed.

Finally, get the transcript in front of someone who uses text alternatives, and let what they miss tell you what you omitted. That review is not a compliance check, and formatting alone is not a test of anything. A descriptive transcript is one companion to one file; it does not replace timed captions, which put the words on screen in time with the speech, or audio description, which supplies the visual account in time with the picture. Those are separate artifacts with separate jobs, and this one is not a substitute for either.

Where the work lands

The finished transcript should read as a small, complete account of one named version of one clip: the speech exactly as recorded and attributed, the sounds that matter, the visual facts the audio leaves out, and the additions marked so a reader can see where the description begins. Where something could not be resolved, say so in the text rather than smoothing it over.

The point is not that a transcript is good practice in general. The point is that this clip makes an argument, and the argument lives in the difference between a sentence about leaving a parcel with reception and the image of someone else carrying it to the elevator. A transcript that captures only the sentence has the words and has lost the pitch.

Frequently asked questions

Why can a caption file be technically correct and still fail a reader of a pitch sample?

It may time and display speech and some non-speech, but it can omit who says the line, that the receptionist does not answer, that a third person takes the parcel, and the onscreen label text. Those omissions can remove the pitch's argument even when no spoken words are missing.

What belongs in a descriptive transcript's header before the timed entries?

The exact filename, duration, and date of the version being transcribed, plus stated conventions for marking speech, description, and text visible on screen or objects. A transcript attached to the wrong cut is worse than none because readers trust it.

How should a transcriber handle a word that cannot be made out?

Write that it is unclear, for example a bracketed [unclear 0:14]. Filling the gap from the approved script, treatment, memory, or automated transcription produces a confident line nobody said.

Why not mention the parcel label as soon as the writer knows what it says?

The transcript should preserve the film's reveal order. In the example, the label is legible at 0:23 and not before; moving it earlier turns a delivery into a warning and drains tension. Place information where the film places it.

Does a descriptive transcript replace timed captions or audio description?

No. It is one companion to one file. Timed captions put words on screen in time with speech, audio description supplies a visual account in time with picture, and the descriptive transcript is not a substitute for either.

More in Advertising Browse all articles