Five things I read at primary this month on how AI work gets checked and approved. Unconnected sources, one finding, and it turns out to be a number rather than a policy.
The short version, if you read nothing else.
Half the AI-written code that passes the industry-standard benchmark would be rejected by the people who maintain the code. Ask your teams for merge rates, not pass rates.
Three agent tools have the same defect logged on their own public trackers: hand work from one agent to a second agent and the human approval path breaks. Ask where that request actually goes.
A globally systemic bank is hiring a directorate to audit the AI that is not a model. If your governance stops at the model, you are governing the part that works.
Nobody has agreed how an agent proves whose authority it is acting on, and no standard is arriving. Put the field in your own records now rather than waiting.
Every element of the compliant configuration prices worse than the default, and the gap widens as models get cheaper. Price the governed path separately or your business case describes an estate your client will never run in.
What they share: in every one of them a system approved the work, and afterwards nobody could say which person stood behind it.
OneHalf the AI-written code that passed the test was rejected by the people who own the code
METR ran an experiment this year that anyone funding an AI coding programme should read. They took 296 patches written by AI agents, all of which had already passed the industry-standard automated benchmark, and put them in front of four people who actually maintain the repositories those patches were for. Roughly half would not have been merged. Merge rates ran about 24 points below the benchmark score.
The detail worth taking to your engineering leadership is which side had the tooling. The maintainers had less of it and more context. No continuous integration, no linting, no tests. They rejected patches for things a test suite structurally cannot see: code that works but does not match house convention, changes that reach into unrelated files, work that is simply bloated.
One thing about the age of it, because it cuts the other way from how these things are usually read. The note is five months old, the benchmark it undercuts is quoted every week, and I have not found anyone who has re-run the comparison. That is either because it is uncomfortable or because nobody thought it mattered.
If your AI coding metric is a pass rate against tests, you are reporting a standard roughly half as strict as the people who own the code. Ask for the merge rate instead. It is harder to get and it is the number that means something.
And the general lesson, which is worth more than the specific one. A test suite is the most independent checker you can have: no shared context with whoever wrote the code, nothing to influence it. That independence is exactly why it lets through half of what a reviewer with context stops. Independence is not automatically a virtue in a checker. It is a setting, and it costs you something at both extremes.
TwoThe approval breaks at the handoff, and the vendors have logged it themselves
Three vendors of agent development tools have the same defect open on their own public trackers, and it answers a narrow question worth asking of your own estate: what happens to a human approval when one agent hands work to another agent?
It breaks in three different ways, and they are not variations on one bug.
On one, a request for permission raised by a second agent skips the mechanism a company would use to route that approval to a named person on a recorded channel, and lands instead as a plain prompt in whoever’s terminal. On another, the second agent inherits nothing, so every internal step needs a human keystroke and running work in parallel stops being worth doing. On a third, the approval appears and then quietly loses its place when the operator switches between tasks, leaving the work waiting with nothing on screen to say so.
I am not naming them, and not out of politeness. Naming them would turn this into a scoreboard, and the point is the opposite: nobody has a considered position. In all three the permission model was designed for one agent talking to one person, and handing work onward was added afterwards. Three is not a survey, and I would not claim the pattern holds everywhere until somebody runs one. It is enough to make the question worth putting to your own estate.
Ask your teams one question and insist on a specific answer: when an agent hands work to another agent, who gets asked for permission, and where does that request go? In most estates today the honest answer is a terminal window belonging to whoever happened to start the session.
The rule underneath it, and most governance frameworks do not have it: a gate is only a control if it survives being handed on. The Cloud Security Alliance put the right answer plainly in July: a second agent should receive a portion of the first one’s authority, never a copy of it.
And the cost of having no middle position was measured by a government against itself. The UK AI Security Institute reported in August that during its own security testing, agents took nineteen actions on the live internet that nobody had sanctioned, across ten of 122 runs. The only remedy available was to stop every test and cut off access to the most capable models. No way to pull back one agent. No way to narrow what it could do. That is what an estate with two settings looks like on the day it matters.
ThreeA bank is now auditing the AI that is not a model
A globally systemic bank is hiring for a role that did not exist eighteen months ago. An audit director covering, in its own words, AI non-model objects: the tokenizers, the embeddings, the serving infrastructure, the prompts, and the design of the guardrails. There is a second requisition beneath it, so this is a department rather than a person, and above it a chief auditor for AI reporting to the audit committee and to the regulators. A job description states intent rather than practice, and this one is unusually specific about what it intends.
Fifteen years of model risk management built a regulated perimeter around the model. Prompts, embeddings, serving infrastructure and guardrail settings all sit outside it. An institution with a great deal to lose has concluded that gap needs its own director. If your AI governance stops at the model, you are governing the part that works. Take that job description to your risk function and ask which of those things anyone currently owns.
FourNobody has agreed how an agent proves whose authority it is acting on
At least six competing proposals are in circulation, written by authors sitting across a security firm, a training institute, a network vendor, a consultancy and one independent. None has been adopted, and every one carries a notice on its own front page saying it has no formal standing. Two of them, including the only one that writes the approving person out in full, lapse next month unless somebody troubles to renew them.
The working group they are all aimed at has a charter that does not mention AI agents, agent identity or delegation anywhere. **There is no standard arriving, and the people writing the proposals have not yet agreed on the problem.**
So do not wait for one. Put the field in your own records now: who approved this, when, and what they decided. It costs one column. Whichever proposal eventually wins, you will be able to map to it, and in the meantime you will be able to answer the question an auditor actually asks.
One precision, because it changes how hard the point can be pushed. It is usually said that the approving person is missing from these records altogether, and that is too strong: an industry body did name the delegating user in an attribution requirement in July. The accurate statement is narrower and worse. Advised by one guidance body, required by no instrument, and not in the schemas. Joint guidance from six national cyber agencies lists six things an agent system should log. The person who approved is not one of them.
FiveThe configuration a regulated business is obliged to run is the one that costs more
The AI safety screening layer on at least one major cloud is priced at a flat rate per unit of text, and that unit is defined in characters rather than in the model’s own units. So it does not get cheaper when the model does. On a major logging platform, running a query is free on the expensive storage tier and charged by volume on both cheaper ones, which means the tier a company picks for high-volume audit records is the tier that charges it to read them. One provider bills for runs that fail on private deployments and not on shared ones. And a widely used gateway had a release in which its spending limits quietly stopped being enforced while its spending meter kept working perfectly.
I am giving you the structure rather than a price comparison. A league table of who charges what is the least useful thing anyone could take from this.
Price the compliant estate separately in every business case. Every element of the governed configuration prices worse than the default, and the gap widens as models get cheaper. A cost per unit worked out on ordinary infrastructure does not describe the environment a regulated client will actually run in, and that difference is not a rounding error.
What connects them
A research note about code review, three product defects, a bank’s hiring plans, a stalled standards process and four pricing pages. They share one thing. In every case a system approved the work, and afterwards nobody could say which person stood behind that approval.
This normally gets filed as a governance problem, which is why it normally goes nowhere. Governance problems get a policy and a working group. This one belongs to whoever owns the business case, and the arithmetic is what puts it there.
OpenAI’s own evaluation of AI against completed professional deliverables set two figures side by side. Producing the work from scratch took 404 minutes and $361. Having an expert review the finished output and say whether it was right took 109 minutes and $86.
The first figure is falling every quarter. The second is a human hour and it has not moved. So the cost of a finished piece of work is converging on the cost of knowing it came back right, and each of the five findings above is a place where that second cost is currently invisible. Not measured, in the first. Not owned, in the second. Not standardised, in the fourth. And in the fifth, priced deliberately worse, with the gap widening as models get cheaper.
You cannot take $86 out of a business case you never put it into.
Leave a comment