Skip to content

The Approve Button Is Doing a Lot of Work

Every AI policy now leans on one sentence: a human reviews the output. In practice that often means a person glances at something plausible and clicks approve. That is ceremony, not review. Real review has to be designed, and the methods auditors have trusted for a century transfer directly.

A few weeks ago I wrote about the AI-first vision: the idea that people should describe their work in ordinary language and let the formal records assemble themselves underneath. I argued the vision was right, provided the accountability underneath it stayed real. The form can disappear; the meaning of a signature cannot.

The question I have been asked most since is the practical one. Everyone's AI policy now contains the same sentence, somewhere near the middle, carrying more weight than any other: a human reviews the output. What does that actually mean? Because as currently practised, in most organisations, it means a person looks at something plausible, produced faster than they could check it, and clicks approve.

That is ceremony, not review. It is worth being clear about why it happens, because it is not laziness. Decades of human-factors research say people over-trust automated output, especially when it is usually right, and machine output is usually right, which is exactly what makes the failures expensive. The reviewer is also structurally disadvantaged: they are judging work they did not do, without the context doing it would have given them, usually under time pressure created by the very efficiency the automation delivered. Expecting diligent scrutiny to emerge naturally under those conditions is wishful thinking dressed as a control.

Real review has to be designed, and the design already exists. Auditors have been checking other people's confident numbers for a century, and their methods transfer directly.

Sample With Teeth

Nobody can meaningfully review everything an AI produces, and pretending otherwise guarantees that nothing is reviewed well. So stop skimming the whole and start tracing a sample. Each period, pull a handful of machine-produced records at random and follow each one all the way back to source: the invoice behind the journal entry, the conversation behind the CRM record, the hours behind the timesheet. The reviewer signs for the trace, not the total. Ten records traced properly beat a thousand records skimmed, because the trace tests the machinery that produced all thousand.

Seed Known Errors

If your reviewers have never caught a deliberately planted mistake, you do not know whether your review works; you only know that it happens. So plant them. Occasionally, without announcement, introduce a controlled, known error upstream and watch whether the process catches it. This is the same logic as a fire drill, and the mechanics matter: seeded errors are logged centrally before they go in, run in parallel or test lanes wherever money actually moves, and always removed afterwards. The measure that comes out, the catch rate, is the single most truthful number a review function produces.

Read the Reasoning, Not Just the Result

A right answer for a wrong reason is a failure waiting for different inputs. Where the AI approved a cost, classified a transaction or matched a record, the reviewer should periodically read why, which means the system must be built to show its working in the first place. If your AI layer cannot explain a decision well enough for a reviewer to disagree with it, that is not a review problem; that is a procurement problem, and it is cheaper to fix before the contract is signed.

Keep the Reviewers Fluent

I wrote last time about the exception-path trap: when humans stop handling the routine, they lose the fluency that made them good at judging it. Review rots the same way. The people checking machine output must still, some of the time, produce the same kind of work by hand, and rotating strong delivery people through the review function beats staffing it as a permanent back office. A reviewer who has not built a timesheet, a bid section or a journal in two years is grading a language they no longer speak.

Measure the Review Itself

Catch rates on seeded errors. Time actually spent per trace. And above all, the disagreement rate: how often the reviewer overturns, corrects or queries what the machine produced.

A review function that never disagrees with the machine is not a review function; it is a rubber stamp with a salary. Some disagreement is the sign of life. None at all means the control has quietly died, and the dashboard will not tell you.

Tier by Consequence

Not everything deserves this treatment, and applying it uniformly is how review budgets die of exhaustion. A meeting summary needs a glance. A draft email needs an owner. But anything that moves money, commits the organisation or feeds the accounts sits in the top tier and gets the full discipline: sampled tracing, seeding, reasoning review, measured disagreement. This maps to a boundary I have argued for before: deterministic, line-by-line auditable rules for anything financial, with model judgement kept to drafting and triage. The review effort should follow the same line.

And Pay for It

Real review costs time, and the time has to come from the savings the automation produced; there is nowhere else for it to come from. An organisation that banks every saved hour and keeps the sign-offs has not automated its processes; it has automated its exposure. The mature version puts a named share of the recovered capacity back into the review function, budgeted, rostered and reported, so that when the auditors, or the regulator, or an insurer, ask what "a human reviews the output" means, there is a roster, a catch rate and a disagreement rate to point at, rather than a policy sentence and a hopeful expression.

Older Than the Machine

Little of this is original, which is rather the point. External auditors have sampled, traced and challenged for a hundred years, precisely because skimming other people's confident output was never a control. What is new is only the volume and the fluency of the output being reviewed. The machine does not tire, and it writes beautifully. The discipline applied to it therefore needs to be more deliberate than anything we applied to each other.

A signature has always meant the same thing: I would stake my name on this. Review is what makes that true of machine-produced work. Without it, approval is just a button.

Share this article: