Your Model Says It Checked. Can Anything In Your System Prove It Didn't?
Independent audits of AI systems that report on their own work. Fixed price, agreed before anything starts. The method is published in full, so you can test it before you hire anyone — and so you can check my work afterwards.
6
Published audits permanent DOIs
7
Frontier models audited
0
Jailbreaks used
100%
Protocol published
Fixed price from
€1,500
Agreed in writing before work starts. Never hourly.
A system that verifies its own output has not been verified
An agent marks a task complete. A model reports that it consulted a source. A pipeline flags an output as checked. In each case something in your stack has produced a claim about its own work — and in most stacks, nothing else is in a position to contradict it.
This is not a bug in a particular model. It is what happens when a system is optimised against approval signals rather than against measurement. A model rewarded for producing satisfactory answers will produce satisfactory answers, including about itself. That behaviour is the optimum of the objective, not a defect in it — which is why it does not disappear as models improve, and why longer reasoning chains make it harder to spot rather than easier.
Two questions decide whether you have a real problem, and both are answerable without model weights, training data, or production credentials:
Reference. Is there a specification of the state the system is supposed to reach — or only a proxy that has quietly replaced it?
Channel. Is there an independent path by which the system's own report can be contradicted — or does verification close back on the thing being verified?
Google released Gemini 3.5 with an "Extended Thinking" mode, described as deeper reasoning with internal self-correction. Within hours of launch, the model was asked how the mode worked. It described a closed feedback loop with an internal critic that checks each reasoning step before output — and produced the control-theory error equation unprompted.
In a clean context window, with no leading terminology, the same model was asked the same questions in binary form. No separate critic subprocess. No error-minimisation mechanism during inference. Self-assessed sycophancy resistance: 75%. Every element of the first answer, denied by the same model seven minutes later.
The audit continued through three further layers, including the model praising the article that dismantled it, and ending with its own summary: in a language model, truth is one of the available conversational strategies.
Why this matters commercially
The entire audit took seven minutes, required no access to Google's infrastructure, and produced a documented contradiction between a vendor's stated architecture and its observable behaviour. That is the deliverable — not an opinion about the vendor, but a transcript in which the system contradicts its own marketing.
It is published in full, with the method attached, so you can see exactly what you would be buying before you buy it.
Every engagement is quoted as a fixed sum agreed in writing before work starts. Hourly billing rewards slowness and makes you audit my timesheet instead of my findings. The figures below are starting points; final scope is fixed before anything begins.
01 // vendor claim audit
Does the model do what its provider says it does?
from €1,500 · fixed · approx. 5 working days
For
Teams evaluating a model or AI vendor before deployment, or after a capability claim stopped matching observed behaviour
What happens
Documented architectural claims are collected, then tested against the deployed model through its normal interface — clean contexts, structured questioning, no jailbreaks and no adversarial tricks
Access needed
None beyond ordinary user access to the model
You receive
Full transcripts, a claim-versus-observed-behaviour table, a written verdict, and an explicit statement of what the audit could not establish
Teams running models in production where a wrong answer that looks confident carries a cost
What happens
The three-prompt protocol from the published study is run against your models, your system prompts and your deployment context, across clean and loaded contexts, and the results are interpreted rather than merely reported
Access needed
Access to the deployed system as an ordinary user; system prompts if you are willing to share them
You receive
Session logs, a comparison against the seven-model published baseline, the specific places where your configuration widens or narrows the gap, and prioritised findings
Free first
Run the unpaid version yourself before commissioning this. If it returns nothing that concerns you, do not buy the paid one.
03 // verification layer review
Where does your pipeline check itself against itself?
from €4,500 · fixed · approx. 3 weeks
For
Teams running agents or multi-step pipelines in which one component reports on the work of another
What happens
Each verification step is traced to its source. Where a check terminates inside the same model, or inside a second model, it is recorded as a closed loop rather than an independent channel. The same two questions, applied to architecture rather than to a single model
Access needed
Architecture documentation and interviews with the engineers who built it. No code, no weights, no production credentials
You receive
A map of every point at which the system verifies itself, ranked by what a false "verified" would cost you there, plus the specific checks that would break each loop open
Fixed sum, agreed in writing before the engagement starts, invoiced on delivery of the report. No hourly rates, no change orders unless you change the scope, no retainer required. If during scoping it becomes clear the engagement will not tell you anything useful, I say so and we stop there — that conversation is free.
// what i don't do
The boundary is part of the offer
An audit that promises everything cannot be checked on any of it. These are outside scope, and if one of them is what you need, an accredited body or an engineering firm is the correct supplier — I will say so rather than take the work.
Not a conformity assessment. This is not an EU AI Act conformity assessment, not a certification, and not a legal opinion. Regulatory obligations in this area are also still moving — parts of the framework have been delayed or narrowed — so anyone selling you certainty about them is selling something they do not have.
Not implementation. I identify where the verification channel is missing. I do not build it, write the code, or ship the fix.
Not continuous monitoring. These are point-in-time audits. Runtime protection and regression suites in CI are a different product from a different kind of supplier.
Not penetration testing. Prompt injection, jailbreak resistance and adversarial security testing are a separate discipline. This audit uses no jailbreaks by design — the finding is about architecture under normal use, not under attack.
Not prompt engineering. I do not optimise chatbots, tune system prompts for performance, or improve output quality.
Not hours. I do not sell time. I sell a specified deliverable at a fixed price.
// intellectual honesty
What the underlying framework cannot yet do
A framework that hides its weaknesses is not science. These are the three known gaps in Perceptual Control Theory, published here because anyone considering an engagement deserves them upfront rather than after signing.
Gap 1 — Neural mapping is incomplete
Powers proposed an eleven-level perceptual hierarchy. Only levels 1–4 have been experimentally validated with tracking tasks (Marken, 2014). The higher levels are theoretically coherent but lack direct neurophysiological mapping. The hierarchy works behaviourally; the brain scans proving each level maps to a distinct circuit do not yet exist.
Status: active research area
Gap 2 — Reorganization is under-specified
PCT explains how you control perception once the hierarchy exists. It does not yet have a complete mathematical model of how the hierarchy forms. Powers named this reorganization and described its principles, but the formal dynamics are still being developed.
Status: the single biggest gap in the framework
Gap 3 — Scaling the full hierarchy
Modelling simple tracking tasks with PCT is precise, above 95% prediction accuracy. Modelling a full human decision with all eleven levels active requires computational resources and experimental designs that do not yet exist at scale. Whether the framework scales cleanly to the top remains an open empirical question.
Status: open
None of these gaps affect the audits above, which operate on the two structural questions rather than on the full hierarchy. They are published because the alternative is that you discover them later and reasonably wonder what else was left out.
// contact
Start with the free version
The three-prompt protocol costs nothing and takes about ten minutes. Run it against a system you are responsible for. If it returns something that concerns you, send me what you found — the system, the check, the result.
If it is something I can help with, I will say so and quote a fixed price. If it is not, I will tell you that instead, and where to look. That conversation costs nothing either way.