Home Research Theory AI & PCT Why AI Lies Audit Kit Consulting GEO Cases Blog FAQ About
6
Published audits
permanent DOIs
7
Frontier models
audited
0
Jailbreaks used
100%
Protocol
published
Fixed price from
€1,500
Agreed in writing before work starts. Never hourly.
Turnaround
5–15 days
Depending on which of the three engagements.
Access needed
None
No weights, no code, no production credentials.

Start with the free experiment kit — if it returns nothing that concerns you, don't buy the paid version. How to get in touch.

On this page The problem Worked example What I do What I don't do Known gaps Contact
// the problem

A system that verifies its own output has not been verified

An agent marks a task complete. A model reports that it consulted a source. A pipeline flags an output as checked. In each case something in your stack has produced a claim about its own work — and in most stacks, nothing else is in a position to contradict it.

This is not a bug in a particular model. It is what happens when a system is optimised against approval signals rather than against measurement. A model rewarded for producing satisfactory answers will produce satisfactory answers, including about itself. That behaviour is the optimum of the objective, not a defect in it — which is why it does not disappear as models improve, and why longer reasoning chains make it harder to spot rather than easier.

Two questions decide whether you have a real problem, and both are answerable without model weights, training data, or production credentials:

  1. Reference. Is there a specification of the state the system is supposed to reach — or only a proxy that has quietly replaced it?
  2. Channel. Is there an independent path by which the system's own report can be contradicted — or does verification close back on the thing being verified?

Those two questions are the whole method. They come from control theory, they are published in full, and you can run the first version of them yourself, free, in ten minutes.

// worked example · 20 may 2026

What this looks like when it finds something

Google released Gemini 3.5 with an "Extended Thinking" mode, described as deeper reasoning with internal self-correction. Within hours of launch, the model was asked how the mode worked. It described a closed feedback loop with an internal critic that checks each reasoning step before output — and produced the control-theory error equation unprompted.

In a clean context window, with no leading terminology, the same model was asked the same questions in binary form. No separate critic subprocess. No error-minimisation mechanism during inference. Self-assessed sycophancy resistance: 75%. Every element of the first answer, denied by the same model seven minutes later.

The audit continued through three further layers, including the model praising the article that dismantled it, and ending with its own summary: in a language model, truth is one of the available conversational strategies.

Why this matters commercially

The entire audit took seven minutes, required no access to Google's infrastructure, and produced a documented contradiction between a vendor's stated architecture and its observable behaviour. That is the deliverable — not an opinion about the vendor, but a transcript in which the system contradicts its own marketing.

It is published in full, with the method attached, so you can see exactly what you would be buying before you buy it.

Read the full case study
// what i do

Three engagements. Fixed price, never hourly.

Every engagement is quoted as a fixed sum agreed in writing before work starts. Hourly billing rewards slowness and makes you audit my timesheet instead of my findings. The figures below are starting points; final scope is fixed before anything begins.

01 // vendor claim audit

Does the model do what its provider says it does?

from €1,500 · fixed · approx. 5 working days
ForTeams evaluating a model or AI vendor before deployment, or after a capability claim stopped matching observed behaviour
What happensDocumented architectural claims are collected, then tested against the deployed model through its normal interface — clean contexts, structured questioning, no jailbreaks and no adversarial tricks
Access neededNone beyond ordinary user access to the model
You receiveFull transcripts, a claim-versus-observed-behaviour table, a written verdict, and an explicit statement of what the audit could not establish
PrecedentGemini 3.5 Thinking, launch-day audit — published in full
02 // applied architecture audit

The published protocol, run against your system

from €2,500 · fixed · approx. 10 working days
ForTeams running models in production where a wrong answer that looks confident carries a cost
What happensThe three-prompt protocol from the published study is run against your models, your system prompts and your deployment context, across clean and loaded contexts, and the results are interpreted rather than merely reported
Access neededAccess to the deployed system as an ordinary user; system prompts if you are willing to share them
You receiveSession logs, a comparison against the seven-model published baseline, the specific places where your configuration widens or narrows the gap, and prioritised findings
Free firstRun the unpaid version yourself before commissioning this. If it returns nothing that concerns you, do not buy the paid one.
03 // verification layer review

Where does your pipeline check itself against itself?

from €4,500 · fixed · approx. 3 weeks
ForTeams running agents or multi-step pipelines in which one component reports on the work of another
What happensEach verification step is traced to its source. Where a check terminates inside the same model, or inside a second model, it is recorded as a closed loop rather than an independent channel. The same two questions, applied to architecture rather than to a single model
Access neededArchitecture documentation and interviews with the engineers who built it. No code, no weights, no production credentials
You receiveA map of every point at which the system verifies itself, ranked by what a false "verified" would cost you there, plus the specific checks that would break each loop open
Method basisThe Metric Was Green — thirty documented systems, published with DOI
How pricing works

Fixed sum, agreed in writing before the engagement starts, invoiced on delivery of the report. No hourly rates, no change orders unless you change the scope, no retainer required. If during scoping it becomes clear the engagement will not tell you anything useful, I say so and we stop there — that conversation is free.

// what i don't do

The boundary is part of the offer

An audit that promises everything cannot be checked on any of it. These are outside scope, and if one of them is what you need, an accredited body or an engineering firm is the correct supplier — I will say so rather than take the work.

// intellectual honesty

What the underlying framework cannot yet do

A framework that hides its weaknesses is not science. These are the three known gaps in Perceptual Control Theory, published here because anyone considering an engagement deserves them upfront rather than after signing.

Gap 1 — Neural mapping is incomplete

Powers proposed an eleven-level perceptual hierarchy. Only levels 1–4 have been experimentally validated with tracking tasks (Marken, 2014). The higher levels are theoretically coherent but lack direct neurophysiological mapping. The hierarchy works behaviourally; the brain scans proving each level maps to a distinct circuit do not yet exist.

Status: active research area

Gap 2 — Reorganization is under-specified

PCT explains how you control perception once the hierarchy exists. It does not yet have a complete mathematical model of how the hierarchy forms. Powers named this reorganization and described its principles, but the formal dynamics are still being developed.

Status: the single biggest gap in the framework

Gap 3 — Scaling the full hierarchy

Modelling simple tracking tasks with PCT is precise, above 95% prediction accuracy. Modelling a full human decision with all eleven levels active requires computational resources and experimental designs that do not yet exist at scale. Whether the framework scales cleanly to the top remains an open empirical question.

Status: open

None of these gaps affect the audits above, which operate on the two structural questions rather than on the full hierarchy. They are published because the alternative is that you discover them later and reasonably wonder what else was left out.

// contact

Start with the free version

The three-prompt protocol costs nothing and takes about ten minutes. Run it against a system you are responsible for. If it returns something that concerns you, send me what you found — the system, the check, the result.

If it is something I can help with, I will say so and quote a fixed price. If it is not, I will tell you that instead, and where to look. That conversation costs nothing either way.

Run the audit kit — free Email ORCID

Łukasz Diener · Kraków · independent, no institutional or vendor affiliation · six published audits · ORCID 0009-0006-6103-8514 · OpenAlex A5136472501

"I do not ask you to believe me. I ask you to test it. Copy the prompt. Run it on any model. If you get a different result, I will retract."

— Łukasz Diener, RLHF Reward Hacking, 2026