A benefits agency deploys agents to triage claims, and the numbers look good: claims processed faster and more consistently than the human caseworkers managed. This is how such deployments are judged — performance measured against a human baseline. Can the agent handle enquiries as accurately as a person? Can it triage faster, at lower cost? Once it consistently outperforms the human on these measures, the deployment is called a success.
This milestone — sometimes called the crossover point — is a useful operational marker. It tells you something real about whether the system is working as intended. But it has a structural limitation that rarely gets named. It measures performance relative to current human practice. It has no external referent. It cannot ask whether current human practice is the right target. It cannot ask whether the task itself should be done differently, or at all. It cannot ask what the organisation would refuse to do regardless of how efficiently an agent could do it.
This is a form of proxy-capture. A measure introduced to track whether the system is doing its job becomes the definition of what the job is. Once the crossover point is reached, the question of what the organisation is for has not been asked. It has been skipped.
The consequence for AI governance is specific. An organisation that defines agent success purely by the crossover point has no mechanism for detecting when its agents are optimising effectively for the wrong thing. If the task being automated is the wrong task, or if the baseline being beaten reflects a practice that was itself flawed, the crossover point will still be reached. The metric will show success. The underlying problem will not appear in the numbers.
Return to the benefits agency. If its agents reach the crossover point — processing claims faster and more consistently than human caseworkers — that is real evidence of operational improvement. It does not tell you whether the criteria being applied are just, whether the edge cases being automated away deserve human attention, or whether faster processing is what claimants actually need. These questions cannot be answered by comparison with the human baseline. They require a prior answer about what the agency exists to do, and what it will not trade away in pursuit of efficiency.
That prior answer has to come from outside the performance measurement system, because the performance measurement system cannot generate it. It can measure against targets. It cannot choose them. Adding more metrics does not close this gap. More accuracy measures, more bias checks, more audit trails are all necessary and none of them supply the missing reference point. They still measure against the baseline already in use. They do not ask whether that baseline reflects the right purpose.
The question that fills the gap is something like: what does this organisation exist to be, and what will it not do regardless of efficiency gains? This is not abstract. It has direct operational content. The answer determines where human oversight is required regardless of how accurately the agent performs. It determines what categories of decision should not be automated irrespective of performance metrics. It determines what counts as a good outcome when the numbers are ambiguous.
The Goodhart’s Law version of this problem is well understood: when a measure becomes a target, it ceases to be a good measure. What is less often named is the version that operates at the level of organisational identity. The crossover point is not just a bad metric. It is a metric that can show green while the organisation is drifting away from its own purpose, because the purpose was never formally part of what the metric measured.
Most current AI governance frameworks do not require an answer to the identity question. They require performance targets, bias assessments, and audit trails. These tell you whether the system is doing the defined thing well. They do not ask whether the defined thing is right. The question of whether the baseline reflects the right purpose is not on the form.
This is frame-failure operating at a level that standard monitoring cannot reach. The frame determines what can be measured. What falls outside the frame is not flagged as absent. It simply does not appear.
Related
- Proxy capture — the pattern where a measure introduced to track something real becomes the target, and the underlying thing quietly degrades.
- Frame failure — what happens when organisational assumptions no longer fit the environment, producing symptoms that more effort inside the frame cannot resolve.
- Purpose, task and the problems nobody has named yet — on the gap between what organisations say they exist to do and what they actually measure.
- Why the frame cannot see itself — why the systems that should detect this problem are typically built from the same assumptions as the problem itself.