Skip to content

The AI Can Act, Not Just Recommend. How Must the Sales Pitch Change?

Business

The AI Can Act, Not Just Recommend. How Must the Sales Pitch Change?

The demo holds up right until the roadmap slide. The assistant has just drafted a customer follow-up — plausible tone, correct facts, the kind of output that closes a meeting. Then someone says the next release will send it for you. Nothing on the screen changed. What changed is the buyer's exposure, and the room can feel it even when no one says so.

Most AI sales material is built to prove output quality, because that's what used to be for sale. Once the product can write to another system, change a record, or deliver a message, output quality becomes one of two things you have to explain. The other one is authority: what this thing is allowed to do, on whose behalf, after what check, and what it costs when it does the wrong one.

So the question isn't "how do I make autonomy sound safe." It's how to describe a delegated action precisely enough that a buyer can decide what to delegate.

Find the first action that leaves the conversation

Start with one task from the buyer's actual work, not a product category. "AI for customer communications" has no boundary to examine. One ticket does.

Walk that task until it changes something outside its own draft — a row in a database, a message in someone's inbox, a status another team reads. That's the line. Drafting, summarizing, ranking, suggesting: all of it stays inside the conversation and costs the buyer a review pass if it's wrong. Everything after that line costs them something they can't get back by pressing delete.

Then name four things about the crossing point:

  • The target. Which system, which record, which mailbox.
  • The operation. Send, update, close, pay, delete, escalate. Be specific about the verb.
  • The initiating authority. Whose permission this action is spending — the user's session, a service account, an integration.
  • The approval point. Where a person sees the thing before it becomes true, and what exactly that person is being shown.

If you can't fill those four lines in for your own product, no slide deck will do it for you. If you can, you already have the most interesting part of the pitch.

What follows is an invented storyboard for a pitch. It is not a test, a demonstration, or a record of anything that happened. It's the kind of example a seller has to construct before the real one exists, and it's here to make the boundary concrete.

Kestrel Industrial Supply, a made-up distributor, runs a service desk. Ticket #4182 is a delayed-shipment complaint from Meridian Fabrication about a replacement pump seal. The complaint gets resolved: the seal shipped, it arrived. An account manager opened the ticket by forwarding the customer's note, so the ticket's Requester field holds the account manager's mailbox, while Meridian's buyer sits on the order record as the delivery contact. The assistant's job is to draft the follow-up.

Keep that field in mind. Nothing in the draft will be wrong. The draft is fine. The destination is the problem, and it produces no error message when it's wrong.

Three ways to hold the decision

Suggestion only. The assistant produces the draft and a proposed status change. A person reads it, checks the address against the order record, sends from their own mailbox, and sets the status. Every consequential decision stays with a human being, which is the honest version of "the assistant only suggests." One thing to say out loud in the pitch: the assistant's own confidence isn't part of the control. If the model is wrong, the reviewer catches it or doesn't, exactly as with any draft a colleague hands over. The buyer gains writing speed. They don't gain sending throughput.

Confirmed action. The operation executes, but only after someone approves it. The load-bearing word here is corresponds. An approval has to correspond to the actual target, the actual content, and the actual operation — not to a summary of them.

Here's the storyboard's first confirmation screen: it shows the message body, and the button says Send to requester. It resolves "requester" to a mailbox the reviewer never sees. The reviewer approves, and the follow-up lands in the account manager's inbox. It reads perfectly plausibly there — it's a shipment update about a ticket that person opened. So nobody bounces it back, nobody asks why, and the activity log records a send to an address that is only wrong if you happen to know the customer contact lives on the order record.

Now the same task with a second confirmation screen: the resolved address displayed as an address, the body, the operation named ("send email to this recipient"), and a separate line for the status change. Same reviewer, same draft, different decision — reject the send. Correct the recipient to the contact on the order record. Approve the corrected version. Note that the corrected approval is a fresh decision about fresh content; approval isn't a license covering a class of sends. The rejection lands in the log with a reason.

The control that saved the day was not the existence of a confirmation step. It was what the confirmation showed. "Human in the loop" describes a diagram. It doesn't describe a control.

Bounded automatic action. The operation runs without a person when named conditions hold. On this storyboard, the conditions a seller would have to state out loud are: the order record shows the shipment delivered; the body is one of the pre-approved variants with the tracking number pulled from the order record; the recipient is the contact on the order record; there's no open escalation flag on the ticket. On #4182 the recipient condition fails, the assistant holds the send, shows the mismatch, and waits.

Notice what the conditions caught: not a subtle judgment failure, but an ordinary field mismatch. That's what conditions are for. They are also only as good as the fields they read. If "delivered" comes from a carrier feed that lags by a day, the condition can be satisfied by stale data, and the buyer deserves to know the staleness window. If you can't name the data sources behind your conditions, they're hopes with names on them.

A held action also has a price. Someone has to work the queue, or a customer who should have heard back doesn't. Name the owner of the exception queue and how long an item is allowed to sit in it. "It escalates to the user" is only an answer if a specific user is watching.

One more thing about presenting all three: fewer confirmations is not a score. A demo that advertises "zero clicks" is selling an authority change as a user-experience improvement, and the buyer is the one who pays the difference. Present each route with the decision it moves and the person who now has to live with it.

What can be stopped, reverted, and lived with

Interruption first, route by route. In suggestion-only, anyone can stop at any point, because nothing has moved. In confirmed action, the reviewer can stop until they press the button — and only if the screen gives them the operative facts. In bounded automatic action, the condition gate is the interruption, plus whoever owns the held queue. Once an operation commits, nobody "stops" it; there's only recovery, and recovery is a different activity performed by a different role.

Then the crux, because decks tend to use one word — undo — for two unrelated things.

Undoing a local record: the status goes from Resolved to Reopened. In this construction, resolving the ticket stopped an SLA clock and queued a satisfaction survey to the Requester address; reopening restarts the clock from the reopen time, so the elapsed time isn't restored, and the survey already queued isn't recalled. The activity log keeps both entries. Reverting a field is not restoring a state.

Undoing a delivered message: harder, and the storyboard's honest answer is that the note is out. The correction is another note. Ask your own product what it means by undo after delivery — if it's a recall, establish what that feature actually guarantees in the buyer's environment, whether it survives the message being read or forwarded, and whether your product even controls it. If your recovery story is a recall request filed under "undo," it belongs under "ask your mail administrator."

Partial completion is the case most pitches never reach. The send succeeds and the status write fails. The customer has the note; the ticket still looks open; the clock is still running. Which system is the source of truth, what does reconciliation look like, and whose job is it? If the answer is "the user notices eventually," that isn't a recovery path.

Route Who can stop it before commit What the record shows after What can't be recovered
Suggestion only anyone, at any point the reviewer's own send nothing new — the human action is the human's
Confirmed action the reviewer, if the screen names the target and the operation reviewer, timestamp, target, content a committed send; a queued survey
Bounded automatic the condition gate; the exception-queue owner condition result, hold reason, or commit anything downstream of a stale field the condition trusted

One field survives this whole analysis: the Requester address, which is still the account manager's. Even after the reviewer corrects the email recipient, the status change still queues the survey to the wrong mailbox — the visible operation got fixed and the data the other operation reads didn't. Fixing a misdelivery and fixing the field that caused it are separate jobs, and someone needs to own the second one.

Match each control claim to implementation evidence

OWASP's LLM06:2025 guidance on excessive agency separates permission and delegated action from output generation, and discusses execution in a user's context, approval for consequential actions, and authorization enforced by downstream systems. That's a bounded reading of a security guidance page (read 2026-09-18; an earlier reading on 2026-09-08 recorded the same scope). It doesn't certify any product. It won't tell you that your confirmation screen displays the right facts, or that your permission boundary is enforced where the consequences land.

The distinction to import is this: your approval interface is a claim, and the enforced boundary lives in the receiving system and in the account the action runs under. Three questions follow from that, and a seller should be able to answer them in the room:

  • Does the assistant act in the user's context, or as a service account? If it's a service account with broad write scope, the user's intent isn't the limit of what's possible, and the confirmation screen is decoration over a wider permission than the reviewer believes they're granting.
  • Does the receiving system enforce its own authorization on the write? If it trusts whatever calls it, your control is the only one, and it had better be the one you described.
  • What does the audit record contain, and can the buyer read it without asking you for a query?

Then propose a bounded evaluation instead of announcing a conclusion. A pilot that uses the buyer's own permission model, a handful of actions that actually matter, the confirmation screens as shipped rather than as configured for the demo, and a walkthrough of the failure path with the customer's data. Ask to see the recovery workflow exercised — and if a path hasn't been exercised, say so in the pitch rather than describing it in the present tense as working.

One clean run in your tenant is a sample of one, under your configuration, on your data. Don't fill the gap with an accuracy percentage, either. Accuracy and tone metrics measure the sentence. They don't measure whether the sentence arrived somewhere it had no business arriving. That's the failure you're insuring against, and no output-quality number speaks to it.

Say the boundary in the pitch

You don't need a security whitepaper in the deck. You need a paragraph, per action, that a buyer could repeat to their own engineering team without adding caveats.

The follow-up assistant writes to [system]. It runs as [identity] with [scope]. It [drafts only / executes after a person approves a screen showing the target, the content, and the operation / executes without a person when these conditions hold and these data sources are current]. If the target is wrong, [role] can stop it before [commit point]. After that, [what reverts] and [what doesn't]. [Role] owns the exception queue and the recovery follow-up.

Six sentences. If the last one is hard to write, that's information about the product, not about your writing.

Put the failure beside the success in the same deck, with the same weight and the same font. The wrong-recipient case is more persuasive than the happy path, because it demonstrates that you know where the edge is — and a buyer who has been sold autonomy before recognizes the difference between a vendor describing a control and a vendor describing a feeling.

Watch three phrases in your own material. "Human in the loop" without saying where in the loop and what that person sees. "The user stays in control" without saying over which operation. "The assistant only suggests" printed a few slides before the roadmap that sends. And retire the product-level claim entirely: "our AI doesn't take actions" isn't a property of a product, it's a property of one action in one workflow under one configuration. The same product can be draft-only in one place and write-capable in another, so that sentence has a shelf life measured in releases.

The boundary the buyer is actually being asked to accept

Go back to #4182, with both outcomes side by side.

In one, a reviewer approves a screen that said Send to requester, and the follow-up reaches an internal colleague who didn't need it. That one approval carried both operations — the send and the write that marked the ticket Resolved — because the screen named only the first. The customer never hears from Kestrel. A survey goes to the wrong mailbox, an SLA clock stops on work that's only half finished, and the whole thing surfaces weeks later in a conversation about whether the assistant is trustworthy.

In the other, the screen names the address, the reviewer rejects the send, corrects the recipient, approves a corrected version, fixes the Requester field, and the customer gets their follow-up. The cost is a review step that takes a few seconds and a condition that held instead of firing.

Those two outcomes are separated by information, not by intelligence. The first reviewer wasn't careless. They approved the thing they were shown.

That's the boundary a buyer is being asked to accept when they sign: a specific role, accountable for a specific class of operations, with a specific set of facts visible at the moment of decision. Everything else in the pitch — the model, the tone, the latency, the integration list — is downstream of whether that sentence is true. And the evidence to judge it isn't a demo. It's the buyer's own configuration, their own permission model, their own failure path, and the one field in the ticket that nobody thought to check.

Frequently asked questions

What changes when an AI can act, not just recommend?

Output quality remains one thing to explain, but authority becomes the other. Once the product can write to another system, change a record, or deliver a message, the buyer needs to know what it is allowed to do, on whose behalf, after what check, and what it costs when it does the wrong one. The demo may look the same, but the buyer's exposure has changed.

What four things should be named at the first action that leaves the conversation?

Name the target — which system, record, or mailbox. Name the operation: send, update, close, pay, delete, escalate. Name the initiating authority: whose permission the action spends, such as a user's session, a service account, or an integration. Name the approval point: where a person sees the thing before it becomes true, and what exactly that person is shown.

Why can “human in the loop” be insufficient as a control description?

It describes a diagram, not a control. An approval has to correspond to the actual target, content, and operation, not to a summary of them. In the storyboard, a screen saying “Send to requester” resolves to a mailbox the reviewer never sees; the reviewer approves the thing they were shown, and the follow-up reaches the wrong inbox. What matters is what the confirmation displays.

What is the difference between undoing a local record and undoing a delivered message?

Reverting a field is not restoring a state. Reopening a ticket restarts the SLA clock from the reopen time, does not recall a survey already queued, and leaves both log entries. Undoing a delivered message is harder: the honest answer in the storyboard is that the note is out, and the correction is another note. If undo means recall, establish what that feature actually guarantees in the buyer's environment.

What should the pitch say about authority, recovery, and evidence?

Give a per-action paragraph: which system it writes to, what identity and scope it runs under, whether it drafts only, executes after approval, or executes automatically under named conditions, where it can be stopped, what reverts and what does not, and who owns the exception queue and recovery follow-up. Match control claims to implementation evidence. A bounded pilot should use the buyer's own permission model, the confirmation screens as shipped, and an exercised failure path.

More in Business Browse all articles