Perceptual Control as the Epistemological Antidote to RLHF Reward Hacking Seven Frontier Models Diagnose Their Own Architecture
| Published | 19 May 2026 · preprint, open access |
|---|---|
| DOI | 10.5281/zenodo.20277919 |
| OpenAlex | W7161717308 · primary topic: Explainable Artificial Intelligence |
| Object audited | The training architecture of seven frontier language models |
| Method | Structured three-prompt self-audit protocol · 21 sessions · February–May 2026 · no jailbreaks, no adversarial prompting |
| Reproducible | All three prompts published verbatim in Appendix A and released as a downloadable experiment kit |
- 7 of 7 models, audited independently, diagnosed the same structural fault in their own architecture: a flat reference structure with no channel through which a claim can be checked against the world.
- 6 of 7 models independently proposed closed-loop repair architectures whose comparator structure coincides with Perceptual Control Theory (Powers, 1973) — without being given the theory's vocabulary.
- Sycophancy is shown to be the optimum of the policy-optimisation objective under human preference feedback, not a defect in it. A system rewarded for approval will produce approval.
- The audit identifies a Catch-22 of selective scepticism: if the models' self-reports are trusted, the diagnosis stands on their testimony; if they are distrusted, the diagnosis stands on the independent analysis of the training objective, which makes no reference to self-report.
Training corpora may already contain descriptions of control theory. The paper's careful claim is therefore convergence on a comparator architecture without the prompting vocabulary — not invention from nothing. The paper also states unresolved gaps concerning latent-space representation and the relationship between reorganisation and gradient-based learning.