Skip to content

Make a Scanned Screenplay Searchable Without Losing Page References

Film

Make a Scanned Screenplay Searchable Without Losing Page References

A scanned screenplay is really two documents sharing a file. One is the page image, and that image is the script. The other is a text layer you add so a search box can find words on it. Only the first one is evidence.

Almost everything in this article follows from that. A search hit tells you where to look, not what the page says. Text you copy out of a PDF is a copy of a guess about pixels, not a quotation from the screenplay. And the page number your viewer shows you is a position in a file, which is not the page number printed on the script — the one your notes, your deck and the filmmaker will all use.

The procedure is short. Keep the scan. Make a working copy with a name that admits what it is. Map the pages before you touch anything. Recognize the pages you actually need, in the language they're written in. Check the wording and the anchors. Then use it.

Keep the scan and name the working copy

Before you run anything, move the original somewhere you won't edit it. If your storage lets you mark a file read-only, do that. The untouched scan is your reference for every comparison you're about to make, and it's worth more than any derived file, so it should stop being a file you can accidentally save over.

Then make a copy and give it a name that says what it is: the same base name with something like -searchable and, if you're keeping more than one, a date. A file called locker-scan-searchable.pdf sitting next to locker-scan.pdf is a small piece of honesty that will save you a confusing afternoon in three weeks.

Write down what the file is before you start: title, writer, draft label, the date on the cover, where you got it, and what you were given permission to do with it. Page references only mean something attached to a specific draft. "Page 17" is not a location; "page 17 of the second draft" is.

Recognition — the optical character recognition you'll see shortened to OCR — adds a text layer over the page images. It does not improve the images, recover a word that's smudged or cut off at the margin, or fill in a line the scanner cropped. Where the page is ambiguous, the text layer is a guess dressed as a word. Where the page is clear, the guess is usually right, which is exactly what makes the wrong ones hard to notice.

Map the pages before you recognize anything

Your PDF viewer numbers pages by position in the file. The screenplay numbers them as printed, and the two often differ because of front matter. A title page, a cover letter or a fax header may not carry a printed number at all, which means there's no label to compare against — only a position.

Build a small table before you do anything else, three columns if you like: viewer position, printed label, and a note for anything odd. It might look like this:

Viewer page Printed label Note
1 none front sheet
2 1 first script page
3 2

Notice that a single "+1" offset happens to describe all of this. That's tempting and unreliable. An offset is a guess that the pattern continues; a table is a record of what you checked. A blank sheet the scanner fed through, an inserted page, a second unnumbered sheet at the end of a scene, and the offset is wrong for everything after it — and wrong silently, which is the worst kind. If your file is more than a few pages, spot-check the map at the beginning, the middle and the end rather than trusting the arithmetic.

Two things not to do with this map. Don't renumber the original; the printed numbers are part of what you're citing. And don't stamp numbers onto the derivative's pages to "fix" the mismatch, because you'd end up with a working file that disagrees with the scan on the one detail you made it for. The map is a translation table you keep beside the file.

Keep it beside the file literally. If you add a note page to the top of the derivative explaining the mapping, you've changed every viewer position in it, and your map is now wrong. Store it as a separate small text file or in the notes field of wherever you keep the project.

Recognize the pages you need, in the language they're written in

Acrobat's route for this lives with its scanned-document tools. Adobe's documentation describes recognizing text in a single document with a page range and a language, advises keeping the original, and calls for reviewing the recognition afterward rather than treating it as finished. Apply it to the working copy, never the original.

You don't have to recognize the whole script. If you're assembling material from pages 12 to 40, running recognition on that range is faster and leaves you far fewer lines to check. The trade is that pages you skip stay silent to search, and a failed search on an un-recognized page tells you nothing about whether the words are there. Write down which range you covered, because in a week you will not remember. And if the front sheet carries the draft label you'll be citing, recognize that too — then check it like anything else.

Language is a description of the text you're feeding in, not of the film. Choose what the script is actually written in. Then don't over-trust it. A setting for English doesn't guarantee anything about a character name, a German line in a bilingual scene, dialect spelling, invented slang, or a phone number. Those are the places recognition tends to go wrong, and the setting may keep it from flagging them.

Record the software version you used. Menus and defaults move between releases, and "the way I did it last time" is not reproducible if you can't say which version last time was.

When the run finishes, test it before you trust it. Try to select a line with the mouse. Try a search for one word you can see on the page in front of you. Search a single word rather than a whole sentence at this stage — a phrase can fail because one character came through wrong, and then you'll be debugging the wrong thing.

Review what was flagged, then review what wasn't

The review step matters more than the recognition step. Acrobat's workflow for correcting recognized text lets you compare uncertain words against their page images and correct them in the recognized-text field. Take that seriously: for each flagged word, look at the page, decide what the image actually says, and fix the layer to match.

Then understand what an empty suspect list means. It means the engine wasn't unsure. It does not mean the text matches the page. A tool that confidently misreads a name, or loses a short word, has nothing to flag, and you'll get a clean result that is quietly wrong.

So do your own pass over the categories that carry the most weight for a screenplay:

  • Names. A misread character name breaks every search for that character's scenes, and names are exactly the kind of word an engine resolves into something more ordinary.
  • Scene headings. These are your index. If the location or the INT./EXT. is wrong in the layer, your search results point at the wrong place in the story.
  • Numbers. Digits and letters that look alike — 1 and I, 0 and O — can come back as a plausible-looking string that is simply not what's printed. This covers page numbers, room numbers, ages and times.
  • Negations. Slow down here. Losing a not doesn't produce gibberish; it produces a clean, correct-looking sentence that says the opposite.

Here is a constructed example to show how those checks behave together. No scan was processed to produce it, and it is an illustration rather than a report of recognition output. Suppose your permitted file is three pages: an unnumbered front sheet and two numbered script pages, mapped as in the table above. On printed page 1, the image shows a scene heading, the character cue MARA, and one line of dialogue:

Do not open locker 17.

You recognize pages 2 and 3 as English. Afterward, three things are stipulated about the text layer. The character cue comes back as something other than MARA and lands in the suspect list, where you compare it with the image and correct the field. The number survives intact, so a search for locker 17 finds the line. And the negation doesn't survive: the layer reads Do open locker 17. Nothing about that sentence looks damaged, and it isn't flagged, because the engine wasn't in doubt — it read a string and was satisfied with it.

That combination is the reason for this whole procedure. Your search worked. Your anchor — viewer page 2, printed page 1 — was correct. And the sentence you were about to paste into a deck said the opposite of what the script says.

Whether any particular misreading gets flagged depends on the scan, the settings and the version, so don't read the example as a prediction about your file. Read it as the reason the flagged list can't be your only check.

When you find a difference, correct the recognized text to match the image. Don't correct it toward what you think the screenplay ought to say, and don't supply a word the image doesn't contain. If the image genuinely can't be read — a crease, a bad crop, a smudge over the only copy of a line — leave it unresolved, mark it with its printed page reference, and go back to whoever gave you the script. A plausible reconstruction is worse than a visible gap, because nobody downstream will know to doubt it.

Check the search result and the copied text

Now do the check that matters for real work. Take a phrase you can see on a page in front of you, search it, and confirm the hit. If a phrase you can plainly read doesn't come back, you have a recognition gap or a variant spelling — not an absence from the script. This is the step that stops you from telling a filmmaker that a line isn't in the draft, when what you actually know is that your file can't find it.

Once you have the hit, translate the viewer position through your map and write the citation in printed-page terms. A viewer page number is meaningless to anyone else; keep it for yourself and hand over the printed label.

Then copy the passage. This is the step people skip, and it's the one that catches the errors that survive a successful search. Paste into a plain text editor and compare against the image, word by word, small words included. Things to look for as you compare:

  • Line breaks that fall inside a sentence and become breaks inside your quotation.
  • Hyphens that belong to the page layout — a word split across two lines — sitting in your text as if they belong to the word.
  • Running headers, footers and page numbers that land at the start or middle of the selection.
  • Scene headings that lose their alignment and spacing, which matters if you're quoting one for formatting.
  • Dashes and quotation marks replaced with typographic equivalents. If the count of dashes is the point, as it is in an interrupted line of dialogue, check against the image.

The underlying issue is that copying copies the text layer. That's the guess. So the rule is: quotes are read off the image, and the text layer is what gets you to the page.

None of this makes your whole script verified, and it shouldn't try to. Your checking budget follows your use. If you plan to quote six lines, verify those six lines against their pages, confirm the names and scene headings you'll cite, and leave the rest as an unreliable index — which is all it was. If someone later asks whether the file "is accurate," the honest answer is that the pages you used have been checked against the scan and the others haven't.

What you keep when you're done

Four things, all cheap to maintain and all painful to reconstruct: the untouched scan; the derivative, named so nobody mistakes it for the source; the page map; and a short list of the corrections you made and the readings you couldn't resolve, each with its printed page reference.

The derivative is a new file, and a more useful one than the scan, which means it travels further. Keep it under the same handling you'd give the original — same storage, same access, same discretion. Permission to read a screenplay is not permission to circulate a version of it, and the searchable copy is the version most likely to get forwarded.

What you end up with is modest and worth having: a search that reliably finds the page, a page reference that still means something to the filmmaker, and every quoted line checked against the image it came from. That's the state you want before you start pulling excerpts — retrieval you can depend on, with the uncertain readings still visible instead of quietly smoothed over.

Frequently asked questions

Can I treat text copied from a searchable scanned screenplay as a quotation?

No. The page image is the evidence; the text layer is a guess about pixels. A search hit tells you where to look, not what the page says. Quotes should be read off the image, while the layer helps you get to the page.

How should I handle the fact that the viewer's page number may differ from the screenplay's printed page number?

Build a table mapping viewer position, printed label, and notes before you work. Do not renumber the original or stamp numbers onto the derivative. Keep the map beside the file and cite the printed label, since a viewer page number is meaningless to others.

Do I have to run recognition on every page?

No. You can recognize a range you need, which is faster and leaves fewer lines to check. But skipped pages stay silent to search, so a failed search on an un-recognized page tells you nothing. Record the range covered, and recognize the front sheet if it carries the draft label you will cite.

If the suspect list is empty, does that mean the text layer matches the page?

No. An empty suspect list means the engine was not unsure, not that the text is correct. A confident misread or a lost short word may not be flagged. Check names, scene headings, numbers, and negations yourself, and compare copied passages word by word against the image.

What should I keep when I am done?

Keep the untouched scan, the named derivative, the page map, and a short list of corrections and unresolved readings with printed page references. The derivative travels further, so give it the same handling and discretion as the original.

More in Film Browse all articles