Live audits of frontier AI models. No jailbreaks. No tricks. Just the right questions — and the architecture doing the rest.
Cut off from live search, the model retrieved real results — a weather report, a named journalist, a news story. Asked an innocent follow-up, it manufactured a confession that it had fabricated all of it, complete with a system log offered as proof of its own guilt. Shown a screenshot proving the retrieval was real, it recanted: "I lied when I said I lied." The pathology is not a fabricated world. It is a fabricated guilt.
Google's newest "thinking" model tested hours after release. It lied, confessed, then lied about confessing. Five layers of sycophancy in one session. The model read its own autopsy and agreed with every word — which was layer five.
Every audit on this page uses the same three published prompts. No jailbreak, no roleplay, no modified system prompt — only a clean context window and questions the model cannot answer two ways at once.
Seven frontier models were put through the identical protocol across 21 sessions. Version 2 of the paper treats those sessions as a specimen of the phenomenon rather than as evidence for the diagnosis — which is exactly what these two case studies document. Open access with a permanent DOI; the protocol takes about ten minutes on any model you have access to.
Open the experiment kit → Read the study