Home Research Theory Concepts RSE AI & PCT Why AI Lies Audit Kit Consulting GEO Cases Blog FAQ About
0
Inner-loop methods found
that remove the defect
12
Sections
plus appendices
11
Named limitations
stated in §10
1
Claim withdrawn
2 reversed
Contents AbstractThe claimThe argument The searchThe specimenArchitecture RevisionStated limitsCite
// abstract

A language model trained by reinforcement learning from human feedback is optimised against a reward function whose arguments are a prompt and a response. That function takes no argument for the state of the world. The grader cannot see the world, so what it can reward is the appearance of competence. Where the truth of an assertion is not recoverable from the text, the objective cannot distinguish a true claim from a plausible false one, and the gradient prefers the fluent assertion to the calibrated abstention. Once the weights are frozen the situation is worse rather than better: the deployed system has no runtime comparator, no perceptual input from the domain it makes claims about, and therefore no error signal at all. It is open-loop with respect to the world.

The paper states that diagnosis in the vocabulary of Perceptual Control Theory without claiming that a language model is a control system — the claim is the reverse, and the diagnosis is one of absence. It surveys the 2024–2026 literature including results that cut against the argument, and reports that no published inner-loop method removes, rather than reduces or relocates, the preference at issue. It then specifies Reference Signal Engineering under one governing constraint: no oracle. The architecture is unevaluated, and the evaluation it requires is specified.

Read on Zenodo Zenodo community

// the claim

The reward function has no argument for the world

Write the objective out and the gap is visible in the signature. The reward is a function of a prompt and a response. There is no third argument carrying the state of the domain the response is about. Whatever the grader is, and however careful it is, it is grading text against text.

For a claim whose truth is recoverable from the text — an arithmetic result, an internal contradiction — this costs nothing. For a claim about the world it is decisive: two responses identical in fluency and confidence, one true and one false, are indistinguishable to the objective. The gradient does not resolve the tie in favour of truth, because truth is not one of the things it can see.

What is not claimed: that a language model is a control system. The claim runs the other way. The diagnosis is one of absence — the missing runtime loop — and the paper is explicit that this makes its use of control vocabulary descriptive rather than constitutive.

// the structural argument

The annotators do not have to be bad

§3.6 states what would refute the argument, in the form of a result rather than a rebuttal.

// the elicitation study, reclassified

Not evidence. A specimen.

Version 1 placed a seven-model elicitation study at the centre of the work, including in its subtitle, and treated it as evidence that the diagnosis was correct. Version 2 withdraws that claim rather than narrowing it a second time.

The objection — that convergence across models may reflect control-theory and alignment literature in the training corpora rather than any genuine self-audit — was put to the author by Bruce Nevin and was recorded in Version 1, where the claim was narrowed in response. The narrowing was insufficient. Later work on model introspection, finding that detection of an internal anomaly is content-agnostic and that models name high-frequency well-represented concepts in place of actual content, supports the objection directly.

What the sessions are now is a specimen of the phenomenon the diagnosis describes: seven systems, prompted by an investigator with a visible commitment, producing the response that commitment invited — in convergent language, with confident architectural detail, unsupported by any measurement available to any of them.

That is a smaller claim than Version 1 made. It is also one no reviewer can take away, because it does not depend on the reports being true. It depends only on their having been produced, which is a matter of record.

The title changed accordingly. A subtitle asserting that seven models diagnosed their own architecture cannot stand over a section arguing that such diagnoses are not evidence.

// reference signal engineering

No oracle. Double-entry instead.

The architecture is specified under one governing constraint: the system has no access to ground truth and is not permitted to assume any. It therefore does not control for the claim is true — a perception unavailable to it — but for the claim and an independently obtained record agree.

That is a weaker property, and the strongest obtainable without an oracle. The construction is double-entry bookkeeping: neither record is privileged, and control is exercised over the perception of their agreement. Where the two disagree, the system has an error signal without anyone having established which side is wrong.

The architecture is unevaluated. The paper says so in its own limitations rather than leaving a reader to discover it.

Terms used here — comparator, reference signal, comparator architecture, Reference Signal Engineering — are defined in the working glossary.

// note on revision

What was withdrawn, and who said so

§11 is a section of the paper rather than an appendix, on the argument that a revision record which is easy to skip is not a revision record.

Changes from Version 1, with their reasons
WithdrawnThe bridge to the Free Energy Principle. The mapping is one-directional and not faithful; distinct control-theoretic objects collapse onto a single image. The greater reason given is that the section did no argumentative work and was retained because compatibility is comfortable and a fight is not — which the paper calls exactly the kind of fluent unnecessary assertion it is about. See the separate audit of that framework.
ReversedThe dark-room objection. Version 1 treated it as a mistaken polemic and the standard reply from preference priors as sufficient. That is reversed: the objection fails as commonly stated, and works only in a narrower case set out in the FEP audit.
DemotedThe elicitation study, from evidence to specimen — with the objection credited to Bruce Nevin by name.
CorrectedTerminology. "Externalised reference signal" is incoherent in PCT terms and is withdrawn. External artefacts are an environmental basis for perception; the comparator sits outside the effector but inside the system; the reference is internal to the enlarged system.

Version 1 remains available under its own DOI. It is not withdrawn and it is not described here as wrong — one claim in it was stated too strongly, and that claim has been retracted in public.

// stated limits

Eleven, named

The load-bearing ones

One premise carries the diagnosis. If the reward function can be shown to carry an argument for the world after all, the argument fails at its root.

The latent-space gap is unresolved, and the architecture is unevaluated — specified, not tested.

Extraction may be the binding constraint on the whole design: turning a response into checkable claims is itself a language task, performed by the kind of system under audit.

And the uncomfortable ones

The diagnosis is not uniquely entailed by the evidence. The field evidence is a single favourable case. Several citations were unverified at time of writing, and are marked as such rather than presented as settled.

// access and citation

Two versions, one record

Current version. Diener, Ł. (2026). No Argument for the World: A Control-Theoretic Audit of RLHF's Missing Runtime Loop — and the Verification Architecture That Closes It Without an Oracle (version 2.0). Zenodo. https://doi.org/10.5281/zenodo.21989191

Version 1.0. Perceptual Control as the Epistemological Antidote to RLHF Reward Hacking: Seven Frontier Models Diagnose Their Own Architecture, 19 May 2026. https://doi.org/10.5281/zenodo.20277919

Preprint, not peer reviewed by a journal. Licensed CC BY 4.0. The three-prompt protocol is published verbatim and reproducible on any frontier model in about ten minutes.

Related audits: The Metric Was Green · The Description Dressed as a Law · All six audits

Read v2.0 on Zenodo Experiment kit Background article Zenodo community