Skip to content

Pitch a Multi-Camera Comedy as a Performance Form, Not a Camera Count

Television

Pitch a Multi-Camera Comedy as a Performance Form, Not a Camera Count

The pitch that works is about a woman's back.

Dee has until ten minutes to seven to get a broken trophy off a shelf in the greenroom before Rosalind turns around and finds it. Rosalind is at the mirror, practicing her acceptance speech, holding the index card with her notes, with no idea that anything on the shelf behind her is wrong. Toby is in the parking lot with the real trophy in a paper bag.

The pitch that doesn't work is four words long: It's a multi-camera comedy.

Those four words aren't false. They're just empty. They name an arrangement of equipment and skip the thing a reader actually needs — what the audience watches, and where the performance is happening while they watch it. A multi-camera comedy may be shot in front of a room of people or not, in one day or five, with an audience track or without one, and none of that follows from the label. Some of those are decisions for later, and for other people. What you can settle on the page, right now, is the exchange: who wants what, who is standing where, and who can see it.

Describe the exchange before naming the setup

Start with two people and what each of them wants from the other in the next three minutes. Dee wants ninety seconds of Rosalind's inattention. Rosalind wants to get through her own speech without crying at a podium in front of two hundred volunteers. Toby wants to hand over a trophy without anyone learning that it was lost for four days. None of these wants can be satisfied at the same time, in that room, and that's the scene. The camera count has nothing to do with it.

Notice what the opening description gives a reader that a format label cannot: a person, an obstacle, a clock, and a third party. They arrive one sentence at a time: the person and the clock in the first, the obstacle in the second, the third party in the third. The last one is the part pitchers skip. Ask who can witness the attempt, and answer it concretely — Dee, Rosalind, Toby, and the audience, each with a different view of the same shelf.

That difference is where the form lives. When a character is interrupted in a shared space, the interruption happens to them, in public, and the reaction that follows has a duration they don't control. Dee can't cut away from her own face while Rosalind decides whether to thank the board president by name. A held reaction — the thing an actor does while someone else keeps talking — is only available if the two performances are happening at the same time in the same place. That's not a technical boast. It's the reason to build the scene this way, and it's the sentence your pitch is missing.

At the same time, notice what the label doesn't buy you. A pitch that promises "the energy of a live audience" has described an atmosphere, not a scene. If the next question — what is that audience looking at? — has no answer beyond "whatever we cut to," you've written a note about editing.

Let the room participate in the joke

This greenroom is invented for this article: no recording of it exists, no rehearsal of it has happened, and no audience has seen it. It's here because a scene is easier to argue about than a principle.

Picture it as a small room with three fixtures that matter. A mirror on one wall with a counter under it. A display shelf on the opposite wall, chest height, where the Golden Broom normally sits before it's carried to the stage. One door, at the far end.

Rosalind stands at the mirror, facing the wall, so she can't see the room directly. She can see it in the mirror — the shelf, the couch, the door. That single fact does most of the work in the scene. It means Dee is always working inside Rosalind's field of view, and the only thing protecting her is Rosalind's attention, which is currently on her own consonants. The mirror also reverses left and right, a small tax Dee pays every time she reads Rosalind's reflection to check whether her eyes are open.

Here is the situation in objects. The Golden Broom — the real award, the engraved one — went missing from the box office on Monday and is currently in Toby's car, packed by accident into a box of old programs. Its understudy was a duplicate the company bought years ago for the lobby case. The duplicate broke this afternoon, when a stagehand backed a flat into the shelf: base split, broom snapped off. That leaves a canvas tote, a broken trophy in it, and, in the prop loft, a household broom sprayed gold with a ribbon tied around the handle. The photographer from the local paper is coming at ten to seven to take Rosalind's picture with the trophy. So the shelf cannot be empty.

Watch the viewer watch this. Dee lifts the broken duplicate off the shelf and lowers it into the tote. She times it to Rosalind's eyes closing. She brings out the gold broom and sets it where the trophy was. Rosalind, mid-sentence, stops, lifts her glasses, and stares at the ceiling while she tries to remember the name of the woman who runs the box office. Dee freezes with the broom in the air. Rosalind puts her glasses back on and continues.

The audience is watching someone prepare a deception while the person it's aimed at stays unaware, and — this is the important part — the audience can see both halves of that irony in a single frame: the shelf behind Rosalind, and Rosalind's face in the mirror, and the door past both of them. That's a lot of story in one shot, and it's available because the room was built to keep those three things in view.

Now the discipline. None of that is a floor plan. The couch, the mini-fridge, the callboard, the poster from last season — all of them are set dressing until they produce an exchange. The mirror earns its place because a character can work behind another character's back, and that's a joke you can run every week with different people. The shelf earns its place because an object on it can be wrong. A wall earns its place when someone has to walk past it to lie to somebody's face. Write that down and cut the rest. A set that's merely recognizable across episodes is a set you're paying for in rent and in exposition.

Compare a held exchange with a cut-dependent version

Run the same event twice.

Version one — the room held open. The scene plays continuously, all three people eventually in the room. Dee's swap happens inside the viewer's attention: they see the lift, they see the freeze when Rosalind looks up, they see the gold broom land on the shelf. Toby's entrance arrives through the one door, in frame, behind Rosalind, so the audience knows he's there before she does. The last beat is her turn. She sweeps the room once and takes in Toby's paper bag, Dee's hand still on the tote, and a spray-painted broom standing where the trophy belongs. Nobody says anything for a second and a half.

Version two — built from separate shots. The scene is Rosalind's. The camera stays with her, close, through the whole speech. Off-screen, there's a wooden clack from the shelf, and she frowns and keeps going. We stay on her. When Toby comes through the door, we finally go wide, and that first wide shot is the reveal: bag, tote, broom. The audience learns about the substitution at the same instant Rosalind is about to.

What the first version buys is anticipation. The viewer knows something Rosalind doesn't and has to sit with it, which converts the scene from a reveal into a held breath. The Dee performance becomes genuinely two-handed — she is playing against Rosalind's timing, not against a plan. What it gives up is surprise. By the time the reveal arrives, the audience has been staring at the broom for two minutes.

What the second version buys is a clean jolt, and a speech that plays uninterrupted, as one performance, in close-up. What it gives up is the near-miss, the freeze, the pleasure of watching someone else's attention be the only thing standing between a secret and a room full of people.

Neither version is funnier by construction. Suspense and surprise are different goods and a pitch should say which one it's after. A version that keeps Rosalind out of the room entirely — the speech alone, the swap happening where no camera ever looks — can be a lovely scene about a woman rehearsing her own gratitude, and in that one the reveal belongs to her rather than to us.

One thing to keep straight: both versions cut. A multi-camera comedy cuts constantly. Nothing here forbids a close-up, and nothing here forbids withholding. The difference is what the viewer was allowed to watch while it was happening, and who holds the timing. In version one, Dee's freeze exists because Rosalind paused; the pause and the freeze are one performance. In version two, that relationship is assembled in an edit from pieces shot at different times.

There's also a practical wrinkle worth noticing before your collaborators do. If the whole point is that the audience never sees the swap, you may be paying to stage a room in order to get a cut scene. You can still do it — you just want to know that's the trade, and to be able to say why the staged version is worth it anyway. Maybe it's the live reaction of a room full of people. Maybe it's that the actors can find the timing together. Both are legitimate. "It's a multi-camera show" is not one of them.

State the proposed performance and audience relationship

Now write the pitch. Two passages, corresponding to what you just decided.

Held-version pitch. The scene plays in the greenroom with all three of them in the room at once. We hold a frame wide enough to keep the shelf, the mirror, and the door in view together, so the swap is never a secret from the audience — only from Rosalind. Dee's movements are timed to Rosalind's speech; the audience watches her wait for the pause. The button is Rosalind's turn: she takes the room in one sweep and finds three people standing inside three different versions of the same fact. Nobody explains it. We go out on her face.

Cut-version pitch. The scene belongs to Rosalind. We stay close on her through the whole speech, and the room happens off-screen behind her — a clack, a shift, a door. She never turns. Toby comes through at the end, and our first wide shot of the night shows the audience what she's about to see when she does turn: a paper bag, a stage manager with her hand still on a canvas tote, and a spray-painted broom where the trophy should be. The turn is hers, and the audience gets there a breath before she does.

Notice that neither passage mentions cameras. That's deliberate. The number of cameras, whether there's a crowd in the room, whether the show shoots five days or one, how long rehearsal runs — none of that is settled by the words "multi-camera comedy," and all of it belongs to the people who'd have to build it. Say what you intend about rhythm, interruption, and where the viewer sits. Mark the rest as open, and bring it to the collaborators who own it.

For what it's worth, this isn't a private theory. In a 2024 episode of the podcast Scriptnotes — episode 653, "Multi-Cam Comedies and WGA Dollars," with a transcript published alongside it — the writer and producer Betsy Thomas describes movement through a scene, playing for an audience, and keeping a space continuous as practical parts of the work. That's one practitioner describing her own choices in conversation. It isn't a rule the form imposes on everyone, and it isn't evidence about budgets, schedules, or whether any particular show has a live crowd. Take it the way you'd take advice from a colleague in a hallway: useful, situated, and worth checking against your own show.

And be careful with the two adjectives that end arguments. "Theatrical" and "cinematic" are descriptions, not verdicts. A scene staged so the audience can watch two performances at once is theatrical in a specific, defensible way. A scene built so a substitution lands as information is cinematic in a specific, defensible way. Using either word as a compliment or an insult tells your reader you haven't decided which experience you're proposing.

Twenty minutes before the awards, in one small room: Dee is on her knees at the shelf with a broken trophy in a canvas tote, Rosalind is at the mirror with her back to the room saying I don't know that I deserve this, and the door is behind both of them, in frame, with the handle turning. The comedy lives in the gap between what the shelf says and what Rosalind believes, and in keeping both of those things visible at once. Any staging that holds mirror, shelf, and door in one field of view will find it. Move the mirror, put the door off-stage, and you can still tell the scene — you'll just be telling a different one, with the dread pulled out and the discovery pushed later. Say which one you're writing, and let the room do the rest.

Frequently asked questions

Why is calling a project 'a multi-camera comedy' not enough for a pitch?

The label names an arrangement of equipment and can be true while saying nothing about what the audience watches or where the performance is happening. A multi-camera comedy may be shot with or without a room of people, in one day or five, with or without an audience track. What can be settled on the page is the exchange: who wants what, who is standing where, and who can see it.

What should a scene description establish before discussing cameras?

Start with two people and what each wants from the other in the next few minutes, then add the obstacle, clock and third party. In the greenroom example, Dee wants Rosalind's inattention, Rosalind wants to get through her speech, and Toby wants to hand over a trophy without revealing it was lost. None can be satisfied at the same time in that room, and that conflict is the scene.

What does a held, continuous version of a scene buy compared with a cut-dependent version?

The held version lets the audience see both halves of the irony in one frame, so they anticipate what Rosalind does not know; the comedy becomes a held breath, and Dee plays against Rosalind's timing. It gives up surprise. The cut-dependent version buys a clean jolt and an uninterrupted close-up performance, but gives up the near-miss and freeze.

Does writing for a multi-camera performance forbid cuts or close-ups?

No. Both the held and cut-dependent versions cut, and a multi-camera comedy cuts constantly. The difference is what the viewer was allowed to watch while it was happening and who holds the timing: in the held version Dee's freeze and Rosalind's pause are one performance; in the cut version that relationship is assembled in an edit.

How should terms like 'theatrical' and 'cinematic' be used in a pitch?

They are descriptions, not verdicts. A scene staged so the audience can watch two performances at once is theatrical in a specific, defensible way; a scene built so a substitution lands as information is cinematic in a specific, defensible way. Using either word as a compliment or insult suggests the writer has not decided which experience is being proposed.

More in Television Browse all articles