Add Captions to a Film Pitch Clip and Deliver the Right Version
Add Captions to a Film Pitch Clip and Deliver the Right Version
A captioned pitch clip is not a transcript laid over the picture. It is a second, time-coded version of the audible information in the film, plus a decision about how the recipient is going to receive it.
Those are two jobs, and they fail separately. You can have perfect words in cues that land half a second late. You can have immaculate timing in a file the recipient's player never loads. The work is finished when a specific cut has a corrected, timed caption track and the person on the other end can actually watch it with the captions on. Everything below is in service of that.
Find the cut first, then the route
Cue times belong to one particular edit. A cue doesn't say "the caretaker speaks"; it says "the caretaker speaks between these two moments in this assembly." Cut four seconds out of the middle, and the picture and sound after that point move four seconds earlier while the cues after it stay where they were — so those cues now land late. This is the single most common reason a caption track "stops working" between revisions: nothing broke, the film moved.
So the first line of your notes should identify the cut — a version number, a date, whatever your project uses — and the second should identify the receiving route. Who watches this, on what?
The route isn't a formality. A reviewer opening a file on a laptop can load a separate caption file if the player supports it. A recipient dropping the clip into a web review platform may only accept a video file with the text already in the picture. A producer forwarding the link onward may end up watching on a phone, in a lobby, with the sound off. You are choosing between two genuinely different artefacts, and the choice is decided by where the file lands, not by which one you find easier to export.
Before all of that, confirm you have a clean master — the uncaptioned picture and audio that everything else is derived from. A burned-in copy is a derivative. It should never become the only surviving version of the cut.
A word on terms, briefly, because the two get used loosely. Captions carry speech and relevant nonspeech sound for people who can't hear the audio or aren't hearing it well; subtitles often carry a translation for people who can. Your pitch clip probably needs the first. Which language the audience should hear, and whether a translation should exist at all, is a filmmaking decision that sits upstream of your cue sheet.
Correct the words before you arrange them
A draft transcript is raw material. It is almost never correct enough to publish as cues, and the errors that matter are not the embarrassing ones.
Take the central case: a character is working a key in a locked door, and an offscreen caretaker says "Not that key." Drop the "not" and you haven't introduced a typo. You've reversed the instruction. The line now tells her to use the key she's holding. Every viewer reading the caption gets the opposite of what the scene says, and the mistake is invisible on a quick read because the sentence still parses.
Small words carry the meaning. "Not," "don't," "isn't," "only," "before" — these are the words that flip a beat, and they're the ones an automatic pass is most likely to swallow or mishear. Check them against the clip with your own ears, not against the transcript you already have.
Then check attribution. The caretaker is offscreen. If you label the line with a name the audience hasn't been given yet — or a role that the film is deliberately withholding — the caption has spent a reveal the film was saving. That's not a captioning decision; it's an authorship decision wearing a captioning costume. When you're unsure, ask the filmmaker how the voice should be named. A neutral label, or no label at all when only one voice is offscreen, is often enough. What you must not do is invent an identity the cut hasn't earned.
Finally, the sounds that carry information. In our example, a latch clicks, and only then does the character react. The click is the cause; her reaction is the evidence that she heard it. If the caption omits the click, a viewer who can't hear it watches someone turn around for no visible reason. If the picture already shows the key scraping into the lock, the scrape announces nothing new — the distinction isn't "which sounds are interesting," it's "which sounds carry something the image doesn't."
Automatic captioning is a reasonable way to get a first draft. It is not a way to get a result. The W3C's guidance on captions is explicit that automatic output needs review for accuracy, and the practical instruction that follows is simple: treat whatever the machine produces as the transcript you were always going to have to correct.
Turn corrected content into timed cues
The correction is content work. The timing is a separate craft, and it's where a good caption track is actually made.
Adobe documents a route for this in Premiere: you generate caption cues from a transcript, and the workflow gives you controls over the caption format, line length, cue duration and the gaps between cues. That's a description of the controls the documentation lists. It is not a promise about how the result reads — automatic alignment is a starting arrangement of text against time, and every boundary still has to be judged.
Two principles do most of the work.
Keep units of meaning together. A cue that ends mid-phrase makes the viewer carry half a thought across a cut in the text. "Not that key" should appear together, in one cue, in one readable hit. Splitting it as "Not that" / "key" doesn't just look clumsy; it delays the one word that carries the meaning.
Cue boundaries should follow the scene, not the stopwatch. This is where a cue that arrives early does real damage. The latch click happens at one moment. If its cue appears slightly before that moment — because the cue needs to be on screen long enough to read, or because a default duration setting rounded it into place — then fast readers learn the outcome of the beat before they hear it. The click was the reveal. The caption has just spent it.
That tension is real and it doesn't have a single answer. A descriptive sound cue can afford a small lead-in; it's giving the viewer context they can hold. A sound that is the beat should start with the sound and let the moment play. Decide which kind you're dealing with, cue by cue. What you shouldn't do is let a default duration setting decide it for you.
Here is the whole shape of it, as an invented example. The clip and the timings below are constructed for this article; they are not the result of a recorded session, and the numbers are placeholders to be read against your own cut.
| # | In–out | On screen | Note |
|---|---|---|---|
| 1 | 0:03.4–0:04.6 | It fits. | The character, at the door |
| 2 | 0:04.4–0:05.8 | CARETAKER (O.S.): Not that key. | Overlaps cue 1 by about a fifth of a second |
| 3 | 0:06.4–0:07.4 | CARETAKER (O.S.): Try the— | Cut off before the end of the word |
| 4 | 0:07.2–0:07.7 | [latch clicks] | The sound the scene turns on; it begins inside cue 3 |
Two cues here need a decision, not a default.
Cue 1 and cue 2 overlap. The character mutters; the caretaker starts before she finishes. You can stack both lines into a single cue with speaker labels, or you can stagger the display times so one line clears before the next appears — accepting that one of them is fractionally early or late. Both are legitimate. What isn't legitimate is leaving two overlapping cues and hoping the player resolves them gracefully, because a player that puts the later cue on top will simply erase the mutter, and you'll never see that happen on your timeline.
Cue 3 is interrupted. The caretaker's "Try the—" is cut off by the latch. The dash is doing real work: it tells the reader the line stopped rather than finished. You might let the cue run through the click so the interruption is visible in the text, or end it sharply and let the sound cue take over. The choice changes how the beat reads. It's a captioning decision with a dramatic consequence, and it belongs to you to make deliberately.
Compare the two deliveries
Now the part that decides whether any of this reaches anyone.
| Separate caption file | Burned in | |
|---|---|---|
| Where the text lives | Alongside the video, as data | Part of the picture |
| What the player must do | Load it, offer it, position it | Nothing |
| Who controls the look | The player | You, at export |
| Can it be turned off | Yes | No |
| Can it be edited later | Yes | Only by re-exporting |
| Can it cover the image | Depends on the player | Yes — check it |
The separate file — often called a sidecar — keeps the text editable and independent. It also hands control of appearance to software you don't own. The position, font and size you arranged in your editing application are not guaranteed to survive into a different player; the W3C guidance notes that support for caption presentation varies across players, and that's the plain fact of it. A cue sheet that sits neatly under the frame in one application can land across a face in another.
Burned-in captions solve the loading problem by removing it. The text is in the picture, it appears everywhere the picture appears, and nobody has to select a track. The cost is total: the text is frozen into that file, it can't be switched off, and it can occupy part of the frame. If the shot you're covering is the lock, or a face, or a hand — and in a pitch clip it often is — a default position at the bottom centre is a guess you haven't checked.
Whichever you pick, remember what it is. The burned-in file can be the version you deliver — for some routes it's the only one that works — but it is a derivative, and it does not replace the clean master. The sidecar is a companion file that can get separated from its video the moment someone forwards the video alone. Keep the clean master, and if there's any doubt about the route, ask the recipient what they can open.
Watch the file they'll watch
Your timeline is not a delivery. It shows you a preview rendered by the same application that authored the cues, which is the one environment guaranteed to be friendly to them.
Play the delivered file, in the actual receiving route, from start to finish. Give particular attention to four places:
The opening. Cue one is the one most likely to be clipped by an export that starts a fraction late, or missed entirely by a viewer who began before the file loaded.
The offscreen exchange. Does the speaker label appear, and does it read the way you intended? A label that's clear on your timeline can be ambiguous on a small screen.
Overlapping speech. Whatever you decided for those two colliding cues, confirm the player did what you expected. This is where you find out.
The final cue. It should come in on time and stay long enough to read. Nothing is more common than a last line that flashes and vanishes because the export is slightly shorter than the sequence.
And check the boring things with equal seriousness: that the correct caption track loaded, that the version you sent is the version you finished, and that the captions are still in sync at the end of the clip rather than only at the start. Drift is cumulative and easy to miss in a short piece.
If the cut changes — even a trim that seems harmless — the timing check is due again. A successful export proves the export ran. It does not prove the captions were included, selected, or in the right place.
What finished looks like
You are done when three things are true. There is a caption track whose words have been checked against the cut, whose small connective words are the right ones, whose speakers are labelled without giving away what the film is withholding, and whose consequential sounds — the latch, not the scrape — are present at the moments they happen. There is a delivery version that matches the route the recipient will actually use, with the image check done in that route and not in the editing application. And there is a clean master behind both of them, so the next revision starts from the cut rather than from a picture with text welded into it.
Get those three aligned and the clip does what a pitch clip is for: it shows someone the film you're describing, in the room they're actually in.
Frequently asked questions
Why do captions fall out of sync after a new cut?
Cue times belong to one particular edit. If footage is trimmed, picture and sound after that point move while later cues stay where they were, so they land late. Nothing necessarily broke; the film moved. Identify the cut in your notes and recheck timing after any change.
How do separate caption files and burned-in captions differ?
A separate sidecar file keeps text as data alongside video; the player must load, offer, and position it, and can usually turn it off or edit it later. Burned-in captions are part of the picture, appear wherever the picture does, cannot be turned off, and can only be edited by re-exporting. Burned-in may be the only workable delivery for some routes, but it is a derivative and does not replace a clean master.
Which sounds should get caption cues?
Captions should carry sounds that convey something the image does not. A latch click that causes a character to react carries information; a key scrape the picture already shows announces nothing new. Decide cue by cue whether the sound is context or the beat itself.
What should I do with overlapping speech or an interrupted line?
Treat them as deliberate decisions, not defaults. For overlap, you can stack lines in one cue with speaker labels or stagger display times, accepting that one line is fractionally early or late; leaving two overlapping cues can let a player erase one. For an interrupted line, the dash shows the line stopped rather than finished, and you can let the cue run through the interruption or end it sharply and let the sound cue take over.
What must be true before a captioned pitch clip is finished?
Three things: the caption track's words have been checked against the cut, including small connective words, speaker labels that do not give away withheld identities, and consequential sounds at their moments. The delivery version must match the route the recipient will actually use, with the image check done in that route. A clean master must sit behind both versions.