The Commercial Will Use a Synthetic Voice. What Performance Are You Commissioning?
The Commercial Will Use a Synthetic Voice. What Performance Are You Commissioning?
There's a sentence that turns up in treatments in one form or another: the spot will use "a warm, natural-sounding synthetic voice, flexible enough to handle any read." Read as a production instruction, it says almost nothing. Warm describes a color. Flexible describes a menu. Neither tells you who decided where the line turns, which word takes the emphasis, or what happens when the client hears it and wants the realization to land a beat later.
Voice identity and vocal performance are separate purchases. Identity is who it sounds like. Performance is what that voice does with these particular words in this scene. You can settle identity with a signature and still have made no decision at all about the performance.
So commission a vocal action. Say who is being addressed, what changes across the line, and which moment carries the change. Then choose the route — a directed recording, a transformation of an authorized take, or a reading newly generated from text — by asking what has to be preserved and what has to move. The route decides who can originate the performance and who can revise it. Everything below is that sentence, taken apart.
Start with a line short enough to direct
Take one line: "You brought it back." It's invented for this article, and no recording, clone, transformation, or generation was made of it. Then give it a circumstance: someone who had stopped expecting to see an object again is being handed it across a table. The movement is surprise into relief. That's the whole scene.
Now decide the performance. The addressee is the person holding the object. Across the four words, several distinct reads are available. The first three stay level and disbelieving, and the last word carries the drop into relief. Or the whole line stays level and the relief arrives after it, in the breath. Or the surprise is in the first word and everything after it is already softening.
Same voice, same four words, three recognizable performances. That's the point of starting here. If your treatment can't distinguish those, it hasn't described a performance, it has described a purchase.
Write down which parts the scene depends on and which parts can float. In this case the scene depends on the line being short and on the relief being late. It does not depend on a particular pitch. The distinction matters later, because it's the difference between a revision that costs one take and a revision that costs a re-cast.
The last thing this section should do is separate a recognizable voice from a permission. A voice description can resemble a specific performer's — that resemblance is a casting note, not an endorsement, and not a right. We'll come back to that at the handoff, because it's where treatments most often overreach.
Where the interpretation comes from in each route
The three routes are usually discussed as if they were three settings on the same machine. They aren't. They place interpretation in different hands and give it different places to live.
A directed recording. A performer in a room, a director with notes, a second take when the first one isn't right. The timing, the breath, the emphasis, the choice of where the relief lands — all of it originates with the performer, shaped by the director, in the session. The revision mechanism is another take, and it happens in real time. This route is the baseline the other two get measured against, which is why it's worth naming costs even when you're not using it: a performer's time, a room, and a schedule.
A transformation of an authorized source take. Start from a performance that already exists and has been cleared for this use. The interpretation in the output is substantially the one in the source, redistributed through whatever the transformation actually lets you change. Two questions have to be answered before anyone relies on this. First, what did the source perform? A take that puts the warmth on "back" and cuts clean contains no late release; you cannot redistribute a decision nobody made. Second, which parts of the source performance sit inside the route's reach? Don't assume a transformation preserves every detail of a performance, and don't assume a fresh reading will reproduce the source's interpretation either. Those are two different claims and neither one is safe.
I have not verified primary documentation for any specific transformation route here, and route capabilities are exactly the kind of thing that changes between providers and between versions. Establish them from the current documentation for the route you're actually planning to book, and then test it.
A reading generated from text. Here you supply words, choose a model and a voice, and work with the settings the interface exposes. Adobe's help page for its Firefly generate-speech-from-text workflow documents text input, model and voice selection, and settings including language, pitch, and speed, with a preview; I'm working from the version of that page currently posted. That's a genuinely useful set of levers, and it establishes the route as its own way of producing a reading. It also establishes only that. Nothing in that documentation says a given setting produces relief, that a pronunciation is correct, or that a voice is cleared for a commercial. Those are things you hear in the output and things a person confirms.
So when a treatment says "synthetic voice" and nothing else, the reader cannot tell which of these three arrangements is being proposed. One of them has an actor in it. One has a source performance with an authorization attached to it. One has a text box. They are not interchangeable, and the differences show up the first time someone asks for a change.
Test a consequential revision, not just a pleasing first take
Here's the direction. In the first approved read, the relief bloomed on "back." Now hold the level through "back" too, and let the relief arrive after the line, on the exhale. The thought lands a beat late.
Run that note against all three routes.
On a directed recording, it's a note. The performer takes it or asks a question about it, the director listens, and the whole phrase gets re-shaped, so you listen to the whole line again rather than a spliced ending. One more take, maybe two. If the performer who gave you the first read is the one the campaign wanted, that's the cheapest route to the new reading — not because recording is cheap, but because the interpretation is already in the room.
On a transformation, the first question isn't how, it's whether. If the source performance ended on the word with a clean cut, the delayed release isn't in the material. What you're asking for isn't a revision of the source; it's a performance the source never gave. That may mean obtaining a new source take under its own authorization — a production decision, not a settings change. If the route's controls do reach timing and pause, then the question moves to listening: a stretched release can read as hesitation rather than relief, and a button labeled with a timing control is not a description of what came out. There's also a question underneath all of it about whether the changed output still falls inside what the source performer agreed to. That question belongs to the agreement and the people who signed it. It does not belong in a treatment as an assumption.
On a generated reading, the levers are the words and the settings. Punctuation, a word repeated, a small pace change, then generate and preview again. Whether the relief actually lands after the line is a listening judgment on the output, made on the render that will ship rather than on the settings panel. And it should be judged the same way every time: generate the reading, listen, decide. If the only way to get the change is to edit a rendered tail in an audio editor, you've invented a fourth operation with its own artifacts and its own review.
Notice what this exercise bought you. The line sounds fine in every route on the first pass. The revision is what separates them. A route that produces one lovely take and cannot take a note is a bad route for a spot that will get notes, which is every spot.
One more thing belongs in this section rather than the last one: authority. Write down who can request the change, who performs it, and who judges the result. If the answer is "the tool is flexible," nobody owns the revision, and the treatment has promised a capability without naming a person.
Carry authorship and language review into the handoff
By now the brief has a route attached to it, and the route has consequences that need to survive the handoff.
Record four things in plain language. The source performance: whose it is, where it came from, and what it was authorized for. The responsible contributors: who directs the performance, who approves the final read. The permitted identity and use route: which voice, cleared for what, through which arrangement. And review responsibility: who listens to the finished output, and against what standard. You're not writing a legal opinion, and this article isn't one. You're making sure the production questions have owners instead of being discovered after delivery.
Treat each language version as its own performance. Words change length when they change language, emphasis moves, and a line that reads as relief in one language may land as impatience in another. A language setting and a speed setting are not a translation judgment. Someone qualified has to listen to the actual output in the actual language and say whether the meaning and the pronunciation survived — and that person needs to hear the finished read, not the plan for it.
Keep draft, approved, and temporary outputs clearly separated. Temp readings used for timing in an edit have a way of becoming the thing everyone remembers, and a version generated for a client review is not the same artifact as the approved master. Name them, date them, and say which is which.
Finally, when something is unverified, write the question into the treatment instead of writing an unconditional promise over it. Not "a warm synthetic voice with the flexibility of a directed performance," but "a text-generated reading, voice to be confirmed against the performer's authorization, directed to hold the line level and release after it, with the final read reviewed by [role]." The second sentence is longer and less exciting. It's also the one a producer can act on.
What the treatment should actually say
A synthetic-voice plan doesn't need much more room than this:
- The action: hold level through "You brought it back"; the relief arrives after the line, in the breath.
- The route: one of the three, named. A directed recording, a transformation of a specified authorized source take, or a reading generated from text.
- The source: for a transformation, which take, and what it was cleared for. For generation, the words and the voice still to be confirmed.
- Revision: who can ask for a change, how the change is made, and what has to be listened to again afterward.
- Review: who hears each language version, and what they're judging.
- Open questions: any permission or output behavior not yet confirmed, written as a question rather than smoothed over.
That's a performance brief with the uncertainty left visible. Choosing a sound takes a minute. This line takes a performance, a revision, and someone listening. That last part is what the treatment is commissioning.
Frequently asked questions
Why is a warm, natural-sounding synthetic voice not enough to commission?
It describes identity and a menu, not performance. It does not say who is being addressed, what changes across the line, which word takes emphasis, where the line turns, or what happens when the client asks for a later realization. Voice identity and vocal performance are separate purchases.
What are the three routes for producing a synthetic-voice performance?
A directed recording, in which a performer and director shape takes in a room; a transformation of an authorized source take, which starts from a cleared performance and redistributes interpretation through the route's reach; and a reading generated from text, where words, model, voice, and settings such as language, pitch, and speed are used with a preview. They place interpretation in different hands.
How can a consequential revision reveal which route is workable?
Take a first read where relief blooms on back, then ask to hold level through back and release after the line. On a directed recording that is a note and the whole phrase is reshaped. On a transformation the first question is whether the source ever performed a delayed release; if the source cut clean, a new authorized source take may be needed. On a generated reading the levers are words, punctuation, pace, and settings, and the judgment is made by listening to the render that will ship; editing a rendered tail creates a fourth operation with its own review.
What must be established before relying on a transformation of an authorized source take?
Ask what the source performance actually performed and which parts of that performance sit inside the route's reach. Do not assume the transformation preserves every detail, and do not assume a fresh reading will reproduce the source's interpretation. Route capabilities change between providers and versions, so confirm them from current documentation and test.
What should a synthetic-voice treatment record for handoff and review?
Record the source performance and what it was authorized for; who directs the performance and who approves the final read; which voice is permitted through which arrangement; and who reviews the finished output and against what standard. Treat each language version as its own performance, keep draft, approved, and temporary outputs separated, and write unverified permission or output behavior as a question rather than an unconditional promise.