Home Research Theory AI & PCT Why AI Lies Audit Kit Consulting Cases Blog FAQ About
30
Documented
systems
2
Failure
signatures
0
Ground truth
required
1
Comparator
architecture
Contents ProblemTwo signaturesComparator BoardsStated limitsWhat is withheldCite
// the problem

The number was correct. The thing was wrong.

Every case in this audit shares one shape: an organisation watched an indicator, the indicator stayed healthy, and the process the indicator stood for failed anyway. Nobody falsified a figure. The figure was accurate. It had simply stopped corresponding to the thing anyone cared about.

The usual explanations are moral — incentives, pressure, culture — and they are not wrong, but they do not tell you where to look before the collapse. This audit treats the situation as a control problem instead, and asks what structural properties are present in every one of the thirty cases while the dashboard is still green.

// two failure signatures

Both are visible before anything breaks

What the audit looks for
Signature 1A goal replaced by a proxy, with no higher-order reference protecting it. The organisation specifies a state it wants, then measures a stand-in for that state. When nothing above the proxy holds the original goal, optimising the proxy becomes indistinguishable from pursuing the goal — until it isn't.
Signature 2A reported indicator never tested against an independent channel. The figure is produced by the same system whose performance it describes, and verified — if at all — inside that system. There is no path by which the report can be contradicted.

Neither signature requires hindsight. Both are properties of the reporting architecture and can be established while everything still looks fine, which is the entire practical point: the audit is diagnostic rather than forensic.

// the architecture

Detection without access to the truth

The natural objection is that you cannot detect a false report without knowing what is true — and if you knew, you would not need the report. The paper's answer is that you do not need ground truth. You need a second channel that is not downstream of the first.

The comparator reads the reported value against that independent channel and raises an alarm on divergence. It does not adjudicate which is correct. It reports that two sources that ought to agree have stopped agreeing — which is the earliest moment at which anyone can act, and it arrives long before the outcome does.

The alarm is specified to bypass the reporting hierarchy. An alarm that travels up the same chain that produced the figure inherits the failure it exists to catch.

// the governance finding

Why boards are routinely the last to know

A board is nominally the holder of the organisation's top-level reference — the specification of what the organisation is for. In practice it perceives the organisation almost entirely through indicators reported by the management it supervises.

That is a structural description, not an accusation. A controller whose only perceptual input is supplied by the process it controls is not supervising that process; it is being informed by it. The paper's position is that this makes the instrument one of governance rather than audit compliance — it examines whether the measuring architecture can be captured, not whether a figure was calculated correctly.

// stated limits

What thirty reconstructions can and cannot show

Illustration, not guarantee

The thirty cases are reconstructed retrospectively. They establish that detection was possible — that both signatures were present and legible before collapse. They do not establish that deployment is easy, and the paper says so in its own words rather than leaving the reader to infer it.

Calibration decides everything

Live performance depends on the divergence threshold and on the false-positive rate. An alarm that fires constantly is switched off, and a comparator that never fires is decoration. The paper states plainly that this is the part that determines whether the architecture works in practice.

// what is published and what is not

The architecture is open. The calibration is not.

The paper draws an explicit line. Published: the two signatures, the comparator architecture, the alarm routing, the thirty cases and their analysis — everything needed to check the argument or to disagree with it.

Retained: the divergence threshold, the aggregation method, and the scoring of channel independence. These are stated as withheld rather than omitted silently, and they are the part that has to be fitted to a live organisation rather than published as a constant.

Readers who want the architecture applied to a specific reporting system will find that under applied work. Readers who only want to check the argument need nothing that is not in the paper.

// access and citation

Open access, permanent identifier

Diener, Ł. (2026). The Metric Was Green: A Perceptual-Control Audit of Metric Gaming Across 30 Systems — and the Comparator Architecture That Detects It (version 1.0). Zenodo. https://doi.org/10.5281/zenodo.21761284

Preprint, not peer reviewed by a journal. Licensed CC BY 4.0. The deposited text incorporates six corrections applied after an adversarial review of the evidence base prior to deposit.

Read on Zenodo Applied work All six audits