The AI Proposal Removes a Step That Currently Catches Errors
The AI Proposal Removes a Step That Currently Catches Errors
The before diagram has five boxes. The after diagram has four. The box that left said "review," and when someone in the room asks what it caught, the answer that comes back is that it slowed things down.
That is a description of what the step cost. It is not a description of what the step did. The second thing is the one the proposal has to carry, and it does not fit in a box.
What follows is an invented workflow with invented people, built to show a comparison method. Nothing in it has been run, measured, or tested.
Establish what the current step actually catches
A fictional outdoor-gear retailer maintains a creative-asset catalog. Campaign teams pull from it. Each entry has a title, a one-line description, a set of tags, and a link to a specific version of an image or video file.
Nadia, the catalog editor, writes entries from three sources: the studio's change note for the asset, the product master data, and the campaign's tag set. Dana, the catalog reviewer, opens each entry, puts the linked image on one side of her screen and the description on the other, and reads. Her checklist has an item that says the description must match the content of the linked version. If it does, the entry publishes to the catalog. If it doesn't, she either returns it to Nadia with a note, fixes a bad link herself when the fix is trivial, or escalates to the studio when the wrong file is in the library.
Now the item that matters. The studio re-shoots the Ridgeline 22 daypack for the spring set, this time in moss green, and uploads it as version 3 of the hero image. The change note for version 3 reads: color pass, tighter crop, matches studio final. That note is accurate. It is also written by someone thinking about image processing, not about product attributes, so it never says the word "green." Meanwhile the product master lists three colorways for the Ridgeline 22, ember red among them, and the campaign tag set contains an ember tag because a different shot in the same campaign uses it. Nadia drafts the description as "Ridgeline 22 shown in ember red."
Every field agrees with every other field, and the description is wrong.
Dana catches this one. The pack is green in the picture and she is looking at the picture. It takes her a few seconds and no special knowledge, which is worth saying plainly, because the proposal's whole case rests on the assumption that her step is friction. Here it is detection.
What she does not catch is just as important. She will miss a shade difference between two greens, a strap or pocket detail that changed between versions, and anything she speed-reads on a long afternoon. She cannot easily verify that version 3 is the crop approved for the channel, because that check runs against a version-number sheet somebody updates by hand, and it is the item in her checklist most often taken on trust. Two reviewers would disagree about some of the same entries.
So the honest first section of the proposal says: this step detects a visible, categorical class of version-content mismatch, by eye, before publication, with a hold attached. It misses a subtler class. It depends on a resource — Dana's afternoon — and it produces inconsistent coverage at the edges. Both the detection and the limits go in writing, because a control described as flawless will be removed the first time someone demonstrates it failing.
Trace the same error through the proposed process
The proposed route generates the description from the change note, the product master, and the tag set, then checks the generated description against those same fields. Disagreement goes to an exception queue. Agreement publishes.
Dana's read is removed.
Follow the ember-red entry. The generator produces a description consistent with the change note, plausible against the product master, and aligned with the tag set. The check compares it to the change note, the product master, and the tag set. Four sources agree with each other. Nothing routes to the exception queue, because there is no exception. The entry publishes, and from that moment the error is in the catalog, in whatever campaign layout pulled it, and in any partner feed that syndicated it.
The distinction to hold onto is that the automated check tests consistency, not truth. That is not a criticism of it. Consistency checking catches a real class of problems cheaply — a description that references a field the metadata doesn't contain, a tag borrowed from a campaign that closed last quarter, a link to a version number that was never uploaded. Those are worth catching. But a coherent wrong description will pass a consistency check every time, and the ember-red entry is a coherent wrong description.
There is a second trap here. The generated sentence reads well. It uses the right nouns, the right product family, the right register. Plausibility is exactly the property that makes a wrong description dangerous in a catalog, and exactly the property a review-free pipeline optimizes for.
And the shorter happy path proves only that the happy path is shorter. You have just replaced what happens after a failure. That is the part that changed, so that is the part the proposal owes the reader.
Compare retaining the check with moving the control
Two routes are actually viable, and they fail differently.
Retain the read where it is. Dana keeps comparing the description to the image before publication. This route catches the visible class, including the ember-red entry, and it keeps the hold in the same place as the detection — the person who notices can stop the thing. What it costs is Dana's queue time, latency between drafting and publishing, and the inconsistency that comes from human reading at volume. What it still misses is the subtle class. Retaining the check is not the same as solving the problem; it is choosing which failures remain.
Move the verification. Remove the pre-publish read and put the comparison later, where someone already has the image and the description in front of them. Tomás builds campaign placements. When he drops the hero image into a layout, the caption sits underneath it. He is the only person in the proposed flow who will look at both objects together, and he will do it for entries that reach him in time. That is genuinely better than nothing on the coverage he has, and it catches the ember-red entry in precisely the situation where it does the most damage.
What the moved route loses is harder to say in a diagram. It loses the hold: Tomás's authority is layout, so he can flag a mismatch but he cannot edit a catalog description or block the catalog publish, and the correction now travels back upstream through two teams instead of being stopped at one desk. It loses timeliness — the detection has moved closer to the moment the wrong image goes out, so the distance between noticing and shipping has shrunk. And it loses coverage. An entry assembled into a layout gets looked at. An entry that goes out in a scheduled email, syndicates to a partner site, or sits unused for six weeks and then gets picked at 4:50 on a Friday does not. The check's scope changes from every published entry is read beside its image to every entry that reaches an assembler in time is read beside its image, and those are not the same sentence.
Replace it with an exception-only route. This is the most attractive option in a proposal, because it keeps the headcount benefit and adds a box labeled oversight. It is also the one that rests on an assumption nobody has stated. An exception-only route needs a justified way to identify exceptions. Here, the flagging signal is disagreement between description and metadata, and the failure class in question is defined by a description that agrees with every field while no field establishes the fact it asserts. The route cannot represent the error. It does not miss the ember-red entry the way an imperfect reviewer misses things; it has no mechanism that could have flagged it.
Making that route real means adding a comparison against the image itself — a model judging whether a sentence matches a picture. That is a specific, interesting, testable claim. It is not something to assert in a diagram, and it is not established by the output looking right.
Every exception-only route also needs a threshold, and the threshold is a decision with an owner. Tuned to catch everything, it produces a queue nobody reads. Tuned to catch only clear cases, it is a review step wearing a different hat. Someone has to set that dial and revisit it, and the tuning is exactly where the invisible class gets excluded.
Account for transferred work and intervention authority
Removing a read moves work somewhere. It rarely deletes it. Four things have to be specified for wherever the work lands, and "human oversight" answers none of them.
Information. The ember-red mismatch is visible only to someone looking at the pixels next to the words. A version number, a confidence score, and a change note are not that. If the person clearing the exception queue cannot see the asset, they cannot make the judgment the queue exists to produce.
Time. A stated budget per item and a capped queue, plus a rule for what happens when the cap is exceeded. Hold the publish, or let it through? That rule, not the flowchart, is the control. A queue with no cap and no drop rule is a queue that gets cleared quickly at the end of the day by someone who has stopped reading carefully.
Skill. Knowing which differences are material. The studio knows that a color pass and a colorway change look similar and mean very different things. A catalog team may not. If the queue clearer cannot tell those apart, the queue is decoration.
Authority. The ability to hold the publish, edit the entry, send it back for a new asset, and stop the feed — with a deputy, because a control that exists only while one person is at their desk has an availability failure built into it.
This is the line between using a system and overseeing one. Clearing exceptions is using the system. Oversight means someone is assigned the queue, resourced for it, and can stop the process without asking permission from the person whose numbers depend on it going out. The NIST AI RMF Playbook's Map guidance (MAP 3.4 and 3.5, in the record I have from 8 September 2026) frames human roles and oversight as questions of actual roles, resources, and responsibilities in context, not as the presence of a reviewer. That framing matches the problem exactly: a named human without time, information, or authority is a label on a diagram.
Two limits on that citation. It is one public, voluntary framework, and it is not independent validation of anything. It supports asking these questions about a proposed control. It does not establish that any control works.
Put the consequence in the main proposal
Lay the two ledgers side by side, in the summary rather than an appendix.
The benefit column has numbers: reviewer hours no longer spent, handoffs removed, latency between draft and publish. The burden column is where the tell lives. If it reads zero, or if it contains the phrase "minimal oversight required," the comparison isn't finished.
Fill it in words where you cannot fill it in numbers, and mark the unknowns as unknowns with a deadline attached. Queue triage, stated as minutes per item times expected volume, or stated as an open question that has an owner and a date. The coverage change, named explicitly: which traffic no longer receives a read beside its image. The unowned failure class: no role in the proposed process is asked to compare a description against an asset. The cost of a late correction across catalog, layout, and syndication. The share of entries that reach an assembler before they go out.
Then say what is actually being removed. Not delay. A read performed by someone who can see the asset, notice a mismatch, and hold the publish. The proposal should state which failures it intends to stop catching, in the same paragraph as the savings, so the decision recipient is choosing between two error profiles rather than between four boxes and five.
Define a bounded test rather than declaring the design safe
The proposal's ending is not a safety claim. It is a plan for finding out, with the exit conditions written down before anything runs.
A later non-sensitive trial on this catalog should reveal a few specific things. Detection on ordinary entries: for a sample of current entries, does the generated description match the linked version, judged by someone who opens the asset? Detection on the class that matters: the metadata-invisible mismatch. That class may be rare enough not to appear in a sample, so construct labeled cases rather than waiting, and report the construction honestly. The rate that actually decides the exception-only route is not overall accuracy. It is how often the system agrees with itself and is wrong — the confident-and-wrong case — because that is the case the queue is structurally unable to hold.
Record false positives too, as a share of the queue. A queue that is mostly noise gets cleared without reading within about a week, and the nominal control survives on paper while the real one is gone.
Measure the queue clearing time against the budget you committed to. For the moved route, measure coverage: what share of published entries reach a person who sees image and text together before they go out? That single figure decides whether the moved check is a control or a sample.
And measure the baseline on the same set. What does Dana's review catch, and what does it miss, on these same entries? Without that, you are comparing the proposal to an idealized reviewer who never speed-reads, and the comparison flatters whichever option you already preferred.
Who evaluates: someone outside the publish-throughput line, the studio or product person who can say what counts as a material difference, and someone with the standing to reject the proposal. Pre-register the observations that would send it back — any confident-and-wrong case in the invisible class, queue time over budget, coverage below what the current review provides.
Three things have to be settled before the box comes out of the diagram. Two of them are questions about the present process: who sees the error, and what they can do about it. The third is the test you are asking to run, and it should be in the proposal as a test, not as a conclusion. The existing check's limits belong in the same document, for the same reason: you cannot defend a control by pretending it was complete, and you should not remove one without knowing what you are giving up.
Frequently asked questions
Why is 'the step slowed things down' not enough to justify removing a review?
That describes what the step cost, not what the step did. The proposal must establish what the current step actually catches, including the error classes it detects and the ones it misses. In the invented catalog example, the reviewer catches an ember-red description because she sees the image next to the text, but she also misses subtle shade differences, small detail changes, and entries she speed-reads. Both the detection and the limits belong in writing.
Why does an automated consistency check miss the ember-red error?
The automated check tests consistency, not truth. The generator uses the change note, product master, and tag set, and the check compares the generated description against those same fields. In the ember-red entry, four sources agree with each other, so nothing routes to the exception queue and the entry publishes. A coherent wrong description passes a consistency check every time. Plausibility is exactly what makes the error dangerous in a catalog.
What is lost by moving verification to a later assembler?
Moving the check to someone who sees the image and caption together can catch the mismatch where it does the most damage. But it loses the hold: that person may be able to flag a problem but cannot edit the catalog description or block the catalog publish. It loses timeliness, because detection moves closer to shipping. It also loses coverage: only entries that reach an assembler in time are read beside their image. An exception-only route is worse for this error class because it needs a justified way to identify exceptions and cannot represent a description that agrees with every field while still being wrong.
What must be specified wherever removed work lands?
Four things: information, time, skill, and authority. Information means the person can see the asset next to the words. Time means a stated budget per item, a capped queue, and a rule for what happens when the cap is exceeded. Skill means knowing which differences are material. Authority means the ability to hold the publish, edit the entry, send it back for a new asset, or stop the feed, with a deputy. 'Human oversight' answers none of these. The NIST AI RMF Playbook Map guidance frames human roles and oversight as actual roles, resources, and responsibilities in context. That is one public, voluntary framework, not independent validation, and it does not establish that any control works.
What should a bounded test measure before removing the box?
Test detection on ordinary entries and on the metadata-invisible mismatch class, constructing labeled cases rather than waiting for a rare error and reporting that construction honestly. For an exception-only route, the decisive rate is how often the system agrees with itself and is wrong, because that is the case the queue cannot hold. Record false positives as a share of the queue, measure queue clearing time against the committed budget, and for a moved check measure what share of published entries reach a person who sees image and text together before they go out. Measure the existing review on the same set. Pre-register the observations that would send the proposal back.