Turn a Factual-Series Paper Edit Into a Source-Linked Video Assembly
Turn a Factual-Series Paper Edit Into a Source-Linked Video Assembly
A paper edit is an argument made of text. A source-linked assembly is the same argument made of recordings, with a route back to each one. Between those two objects sits a step that text cannot perform: deciding whether the recordings, placed in order, actually support what the paper edit says they do.
That step fails quietly. A selection reads well on the page and arrives in the timeline under the tail of someone else's sentence, or against a face turned the wrong way, or with no picture at all. None of this is visible in a transcript, because a transcript has already decided that speech comes in lines — one speaker, then the other — when the recording may have two people talking at once.
The route below is one way through: inventory the recordings against the paper edit, check the selected words against the audio, build the sequence while keeping two different transcripts straight, look and listen at every join that matters, and hand off a project in which the gaps are still visible.
The exercise this piece works through
Nothing was recorded or assembled for this article, so the example is constructed: two interview recordings and a paper edit that have never existed as files. Its facts stay fixed throughout, because the point is to follow a single sentence back to its source.
The series is about a volunteer-run bicycle co-op. One interview, one subject: Elena, a volunteer mechanic, seated in the shop's back room with the camera on sticks. Two recordings came out of the day, labeled A012 and B004.
In A012 an off-camera interviewer asks, "What happens if someone can't pay for a repair?" Elena answers, "We don't turn anyone away." She continues: "If you can't pay, you sweep the floor, or you come back Tuesday and help with a build." Her first word lands under the interviewer's last one — she begins "We" roughly half a second before "pay?" has finished. In B004, later in the same interview, she says, "Everyone here has something to trade." Between the two recordings the camera moved to the other side of the room: in A012 Elena is angled toward the right of the frame, in B004 toward the left.
The paper edit proposes three beats:
- ELENA (A012): "We don't turn anyone away."
- ELENA (B004): "Everyone here has something to trade."
- NARRATION: "That principle shapes everything, from the parts bins to the schedule."
On the page that is a clean argument in three lines. In the timeline it becomes a specific set of problems, none of which the page could show you.
Match the paper edit to the recordings first
Every selection in a paper edit should name a source and a position you can recover, so the first pass is mechanical: can you open the file, find the words, and hear them?
The snags are rarely dramatic. Filenames don't line up across the camera card, the sound recorder and the transcript. A transcript made against a compressed copy carries a small timecode offset, so the position recorded in the paper edit lands a few seconds from the sentence it names. An answer exists in two takes and the paper edit cites the one you don't have open. None of these announce themselves as errors; they show up as a selection pointing at the wrong sentence.
When you can't locate a passage, mark it missing. Don't reach for a nearby line that says something similar. The two are indistinguishable on the page and entirely different in the assembly, and telling them apart is most of the work.
Import the originals rather than a transcoded or exported copy someone passed along, and keep the relationship between those files and the project recoverable. If what you've been handed is a render, you have a preview, not sources.
Both recordings in the exercise open, and both sentences sit where the paper edit says they do. Everything that follows is still a problem.
Check the source text against the recording
A source transcript is a transcript of the original recordings. It does two useful jobs: it lets you search, and it gives you a place to point. It does not establish that the words are right, that the speaker labels are right, or that the line breaks you selected correspond to anything in the recording.
Two failures are worth listening for specifically.
Speaker labels. Automatic transcription mislabels speakers routinely, and in factual work that is not a typo — it's an attribution error. In a two-person interview, a sentence reassigned to the wrong person changes who is accountable for saying it. Once your sequence puts a claim in someone's mouth, that claim is theirs on screen.
Qualifications. The condition on a statement often sits in the clause immediately before or after the words you selected. A line that reads as absolute in a transcript can turn out to be a line about one job, one year, or one exception.
In the exercise, the first selection is followed in the same breath by its terms: sweep the floor, come back on Tuesday. Read alone, "We don't turn anyone away" sounds like an unconditional welcome. Read with the next clause, it describes a trade — a policy with conditions attached. Which of those the sequence says is an editorial decision, and it is also a decision about accuracy.
The question matters as much as the answer. The paper edit lists only Elena's line; the recording contains the question before it. The transcript draws those as two consecutive lines, one above the other, as though a pause sat between them. It doesn't.
Build the sequence, and know which transcript you are reading
Assemble the identified selections into a new sequence. The moment you do, you have a second transcript, and the difference between the two is where this workflow tends to trip people.
A source transcript describes the recordings. A sequence transcript describes what you built. The two can look almost identical — the same words, mostly the same order — while answering different questions. The first tells you what was said. The second tells you what your assembly currently contains, including every decision you have just made.
Adobe's "Edit sequences using Text-Based Editing" page (last updated January 7, 2026; the published page text was checked in September 2026 for this piece) documents that text edits made in a sequence transcript can move or remove the associated clips in the timeline, and that ripple deletion is part of that behavior. That is a description of documented editing behavior. It is not a test of the software on anyone's footage, and it says nothing about whether a transcript is accurate or a cut is fair.
The practical meaning is blunt. In a sequence transcript, deleting a line is deleting picture and sound; this is not a proofreading operation. If you remove a repeated word because it reads badly, you have changed the cut. What the page documents is that one link: delete transcript text and the clips attached to those words go with it, or the gap ripples closed. The rest of the work — setting frame-level in-points, keeping handles, trimming and balancing audio — happens in the timeline, as a step that follows the text edit rather than something the text edit performs for you.
So before you change a word, confirm which transcript you're looking at. The documented timeline-changing behavior belongs to the sequence transcript; make sure that's the one in front of you.
Then look at the sequence itself. A ripple deletion closes a gap, and closing a gap creates a join. Every text edit manufactures at least one new join, which means it manufactures at least one new thing you have not yet inspected.
Look and listen through every consequential join
A join is where two recordings meet. It is the only place where a paper edit's claims become testable, and it is the reason a text-only approval is worth so little. At a join, check the overlap, the breath, the eyeline, the action, and how one delivery sits against the next.
Handles are the recorded material you keep on either side of a selection — the unused part of the clip at each end — so that you can move a cut point without going back to the source. Normally you keep a few seconds above and below every selection.
Here, you can't. The head of the first selection has no clean handle in front of it, because the interviewer's question is there, overlapping. There is material after the answer ends and none before it begins.
That leaves two honest choices. Keep the question audible, and the assembly now contains a voice the paper edit never mentioned — which is a real decision about how the series sounds, and if you want that question on screen you need a shot of it, which this shoot doesn't have. Or start later inside the answer, where there is air.
The second choice costs you the line the paper edit liked best. "We don't turn anyone away" isn't available with a clean front edge; the cut has to start somewhere inside the answer. That is the ordinary arithmetic of an overlap: the cleanest cut and the best sentence are not always the same sentence.
Now the join between the two lines. On the page they run on from each other, because both are about the same principle. In the timeline, Elena finishes a thought facing right and begins the next one facing left. Cut together, she appears to be addressing two different people from the same seat. The words are continuous; the picture is not. Her delivery doesn't match either — she lands the first line slowly and the second quickly — so the second line reads as an interruption of a mood the sequence has just established.
Your options are limited, and none of them is a disguise:
- Separate the lines. If the recordings hold a shot that can sit between them — a different speaker, hands at work, a wide of the shop — the change of direction reads as a new passage rather than a broken one. Whether that shot exists is a question you answered during the inventory, not one you hope about later.
- Accept the change and say so. Some series cut between differently angled shots as a matter of house style, and viewers read the switch as normal. That's a legitimate choice, but it should be a choice, and it belongs in the notes.
- Drop one of the lines, or find the same idea inside a single continuous take, if one exists.
- Don't hide it. A dissolve or a music hit over a broken join asserts that something happened which didn't. If the words are meant to be continuous, a dissolve contradicts them. If time really did pass, a dissolve is honest — and then the note should say so.
One available outcome is that the two-beat argument in the paper edit isn't supported by these recordings as placed. That's a finding, not a failure. It is what you were building the assembly to discover.
And when you repair a join, you create another one at the new in-point. Look at that one too.
Hand off with the gaps still visible
A factual assembly holds three kinds of material, and they must not look alike:
- Recordings you have, each linked to a source file and a position in it.
- Lines proposed but not recorded — typically narration.
- Images proposed but not shot.
Mark each as what it is, on the timeline and in a note that says what would fill it. The exercise's third beat carries two markers: a narration line that exists only as writing, and the image for it, which nobody shot. A placeholder card stating what should be there is fine, as long as it doesn't read as footage.
Then build the source-selection map: a list, kept outside the project as well as inside it, connecting each moment in the sequence to a source file and a source position. Sequence in-point, source file, source in-point, and what the selection is. This is the "source-linked" part. Without it you have a sequence; with it, a reviewer can go from any moment in the assembly back to the recording and hear it in context — which is the only way anyone can check whether you trimmed a statement into a different one.
Export a viewable review version so people can watch the assembly without opening the project, and be clear about what that file can and cannot do. A render answers "does this work?" It cannot answer "where did this come from?" Only the project and the map can.
Before calling it handed off, open the saved project again. Confirm that the media links resolve, the timing is what you think it is, and the markers are where you put them. If someone else will edit from it, they should be able to open it too. A handoff is only as good as its second opening.
Two things a source link does not do
A recoverable route back to a recording proves where a sentence came from. It proves nothing else.
It is not permission. Being able to point at the exact source position doesn't establish that the person agreed to be used this way, or that the footage is cleared for this series. Those are separate questions with separate answers, and a well-built assembly can be a perfectly organized file of material you may not air.
It is also not a fairness check. A sequence can be scrupulously source-linked and still change what a speaker meant, by trimming the qualification, dropping the question, or placing the answer after a different one. That is why the map and the assembled joins need a reader who is thinking about meaning, not only about whether the cut is clean.
Put the assembly next to the paper edit
The last useful move is the plainest. Set the sequence beside the paper edit and write down the differences: lines that arrived with a question attached, sentences that couldn't be cut cleanly at their start, joins that changed the sense of a statement, material that doesn't exist and was assumed.
That list is more valuable than either document. The paper edit is the proposal; the assembly is what the recordings will actually support; the list is the space between them, and every item on it is now a decision someone can make rather than an assumption nobody checked.
What exists at the end of the exercise is smaller than what the page promised: one answer, possibly without its opening sentence; a second line now separated or flagged; a marker where the narration should be; one join that a person has to decide about. That is a weaker argument than the paper edit made — and it is the first version of it that anyone can evaluate.
End with a sequence that reopens, an export that plays, and a path back to each source. The path is what makes the weaker version arguable: someone can disagree with the assembly and show you exactly where, in the recording, the disagreement starts.
Frequently asked questions
What's the difference between a source transcript and a sequence transcript?
A source transcript describes the original recordings. It is useful for searching and for pointing, but it doesn't establish that the words, the speaker labels or the line breaks are right. A sequence transcript describes what you built and reflects every cut decision you have just made. The two can look nearly identical in wording while answering different questions.
If I delete a line of text in a sequence transcript, what happens to the timeline?
Adobe's "Edit sequences using Text-Based Editing" page (last updated January 7, 2026; the published page text was checked in September 2026 for this piece) documents that text edits made in a sequence transcript can move or remove the associated clips, and that ripple deletion is part of that behavior. That is documented editing behavior, not a test of the software on anyone's footage, and it says nothing about whether a transcript is accurate or a cut is fair. Deleting a line is deleting picture and sound, not proofreading.
Why is the overlapping question a problem for the first selection?
Elena's first word lands about half a second before the interviewer's "pay?" has finished, so there is no clean handle — no unused recorded material — in front of "We don't turn anyone away." You either keep the question audible, which adds a voice the paper edit never mentioned, or start later inside the answer, which costs you the line the paper edit liked best. The cleanest cut and the best sentence are not always the same sentence.
Does a source-linked assembly prove the material is cleared to air, or that the edit is fair?
Neither. A recoverable route back to a recording proves where a sentence came from and nothing else — not consent, not clearances, not fairness. A sequence can be scrupulously source-linked and still change what a speaker meant, by trimming a qualification, dropping the question, or placing an answer after a different one.
What should a handoff contain?
Recordings linked to a source file and a position in it; clear markers for lines proposed but not recorded (typically narration) and images proposed but not shot; and a source-selection map listing sequence in-point, source file, source in-point and what the selection is, kept outside the project as well as inside. Export a viewable review version, and be clear that a render answers "does this work?" but cannot answer "where did this come from?" Then open the saved project again and confirm the media links resolve.