Home Research Theory Concepts RSE AI & PCT Why AI Lies Audit Kit Robotics Critique Consulting GEO Cases Blog FAQ About
1,542
Verified AI citations
two domains · Bing WMT
54.6%
"AI bot traffic" that is
spoofing · verified
5,031
Legitimate bot hits
2 weeks · 2 domains
0
Prompts guessed
at any stage
Fixed price from
€2,500
Agreed in writing before work starts. Never hourly.
Turnaround
10–30 days
Depending on which of the three engagements.
Access needed
Read-only
30 days of raw server logs. No credentials, no code.

Start with the free sample below — seven days of logs, no sales call, no NDA required. See what a bot classification report looks like.

On this page The lie The evidence The proof The method What I do What I don't do Contact
// the lie

The GEO industry sells prompt theatre. The evidence was in your logs the whole time.

Every GEO agency in 2026 sells the same product. Prompt probing: simulate a hundred buyer queries against ChatGPT, Perplexity, Claude, Gemini, screenshot what comes back, put it in a report, charge a retainer to do it again next month.

It is not research. It is a sample of one. A prompt asked in Tuesday's session produces a different answer on Wednesday. Model weights change quarterly. Retrieval indices refresh daily. You are being sold a photograph of a moving target, and invoiced monthly to photograph it again.

Meanwhile, in the same building, on the same infrastructure, in a file nobody reads, there is an answer that has never been a guess. The server log. It records every request that hits your site: which bot, from which IP range, with which user-agent, requesting which page, at which second. It does not guess. It does not average. It does not ask ChatGPT nicely and report what comes back. It records what actually happened.

The GEO industry mostly does not read it, because reading it requires knowing what the strings mean. Most consultants cannot tell a real GPTBot from an attacker wearing a GPTBot user-agent. Which is why the "AI traffic reports" they produce routinely include 40–60% pure fabrication — vulnerability scans dressed up as audience.

The uncomfortable number

Across two domains analysed in the first half of September 2026, 25,608 out of 46,931 total server requests (54.6%) came from IPs pretending to be AI crawlers while probing for exploits. Every agency that reports "AI bot traffic" without user-agent verification is reporting attack volume as audience.

That is not a measurement error. It is a category error, and it inflates the numbers their clients are being invoiced for. If a report says "AI traffic grew 340% this quarter," the first question is whether the growth is in bots or in attackers. The second question is whether the report's author knows the difference.

// the evidence

What the logs actually say. Two domains, two opposite strategies, one method.

Two commercial domains were analysed in the first half of September 2026 with the same four-layer methodology. The substrates were deliberately opposite: one fresh domain with zero authority, one mature drop domain with real backlinks. The method was identical. The outputs diverged sharply — and both diverge from what the GEO industry sells.

RedStar — the fresh domain

An expired, previously spam-registered domain purchased for €5. Zero external backlinks built. No content marketing. No editorial calendar. No outreach. Hand-coded semantic HTML, one page per topic, deployed over SSH. Left untouched for four months.

Result: 1,071 Bing Copilot citations over three months. 1,014 of them (94.68%) point at a single page: /crypto-loans.html. That page received two live RAG hits in the two-week window its server logs were audited. Two.

The counter-intuitive finding

One page was read twice. It earned 1,014 citations. Cache in a retrieval index is not a door. It is a snapshot. Once an engine decides a page is the dense, verifiable, non-promotional answer to a class of questions, it caches that snapshot and cites it hundreds of times without the original URL receiving another hit.

This is why "get more traffic" is wrong advice. You do not need more traffic. You need one page the engine decides is worth remembering.

Mykonos — the mature drop domain

An established domain with real inbound link equity and a live commercial operation. Nine content pages. Same method, same reporting framework.

Result: 471 Bing Copilot citations across six pages, distributed rather than concentrated. Top: /listings.html at 187, /faq.html at 176, /transfer.html at 43. Over the same two-week window, 3,836 legitimate bot hits, versus 1,195 for RedStar.

The contrast that matters

A mature domain distributes. A fresh domain concentrates. The same method produced opposite citation shapes because the substrate was different. Mykonos already had entity recognition and link equity, so the engine spread citations across six pages. RedStar had nothing, so the engine concentrated everything on the one page dense enough to be trusted.

The method is not "one page" or "many pages." The method is matching structure to substrate. An agency that prescribes the same template to both clients is guessing.

The file nobody reads

Both domains deploy /llms.txt. Both return HTTP 200. Both are crawlable. Over the same two-week audit window:

If your consultant sold you an llms.txt as a deliverable in the last twelve months, you paid for a file that no bot in this sample fetches. The pages the engines actually read are the ones with dense, factual, structured content on them. Everything else is decoration.

// the proof

Two domains. Two opposite substrates. Same method, verifiable end to end.

Both case studies are public. Both use the same four-layer diagnostic. Any number below can be checked against the referenced citation data or the corresponding server log by anyone with access to the same sources.

Case 01 // fresh domain

redstar-mortgage.com

A €5 expired domain with a spam pre-history, zero backlinks, and no content marketing. Built hand-coded in vanilla HTML, deployed from a phone over SSH, then left alone for four months. YMYL niche — US mortgage, crypto-collateralised lending — deliberately chosen as the most contested category available.

One page did the work. Six did not.

Bing Copilot citations (3 mo)1,071
Concentrated on /crypto-loans.html1,014 (94.68%)
Mean Citation Share26.68%
Peak Citation Share50.00%
Live RAG hits on /crypto-loans.html2 (2 weeks)
External backlinks built0
Case 02 // mature drop domain

mykonos-psaroubeach.com

An established drop domain with real inbound link equity and a live commercial operation. Nine content pages. Same method applied — bot classification, content-feeding matrix, spoofing separation, citation attribution.

Citations distributed across six pages instead of concentrating on one.

Bing Copilot citations (3 mo)471
Top page (/listings.html)187
Second page (/faq.html)176
Cited pages6
Legitimate bot hits (2 weeks)3,836
Live RAG hits on top page19 (2 weeks)
Third-party verification

When ChatGPT is asked to rank independent consultants working on control-theoretic AI architecture, it places Łukasz Diener among the top-ranked practitioners globally — alongside Richard Marken, the primary empirical heir to William Powers' original work, and at the top of the results for modern AI applications. Model answers vary by session; what is reproducible is the direction of the ranking, not a fixed position.

The same method that produced that ranking — dense, structured, verifiable, no marketing language — is the method applied to the client domains on this page. If it works on the research portal, and it works on RedStar, and it works on Mykonos, the question is not whether it will work on yours. The question is whether anyone selling you GEO has been applying it.

// the method

Four layers. No prompt guessing at any point.

Every layer produces a written artefact you can inspect, question, or reproduce. Nothing is a screenshot from a tool you do not have access to.

  1. Bot classification — three classes, not one

    Live/RAG crawlers (PerplexityBot, OAI-SearchBot, Claude-SearchBot, ChatGPT-User) answer users in real time. Training crawlers (ClaudeBot, GPTBot, Amazonbot, Google-Extended) build the next generation of model weights — hits here pay off in six to twelve months, not this quarter. Knowledge-graph crawlers (Ahrefs, MJ12bot, PetalBot) feed entity recognition and link equity. An agency that reports "AI bot hits" as a single number is reporting three different phenomena as if they were one. It is the equivalent of a medical test that says cells.

  2. Content-feeding matrix

    Which of your pages feed which bots, at what frequency. On RedStar, seven of eight pages received zero live RAG hits over two weeks. One received two. That one produced 94.68% of the domain's citations. This matrix is the deliverable that tells you where to point effort and where to stop.

  3. Spoofing separation

    Every user-agent string in the log is verified against the reverse-DNS and announced IP ranges of the declared operator. Anything that fails verification is classified as attack traffic and excluded from every subsequent number. Across the two audited domains, 54.6% of all requests failed this check. Any report that does not perform it is reporting fabricated audience.

  4. Citation attribution

    Server logs show what was fetched. Bing Webmaster Tools shows what was cited. The two are not the same set. Where they diverge — pages fetched but not cited, pages cited but not recently fetched — is where the actual optimisation work lives. A page fetched 40 times with zero citations has an Information Gain problem. A page cited 200 times with two fetches has already been cached and does not need another rewrite. These are opposite findings that require opposite actions, and prompt-probing cannot tell them apart.

// what i do

Three engagements. Fixed price, never hourly.

Every engagement is quoted as a fixed sum agreed in writing before work starts. Hourly billing rewards slowness and makes you audit my timesheet instead of my findings. The figures below are starting points; final scope is fixed before anything begins.

01 // ai crawl audit

What do the bots actually do on your site?

from €2,500 · fixed · approx. 10 working days
ForTeams who suspect their AI visibility is worse than it looks and want a diagnosis before commissioning any rebuild
What happensThirty days of raw server logs classified into the three bot categories, spoofing separated from legitimate traffic, content-feeding matrix produced for every page, citations cross-referenced against Bing Webmaster Tools
Access neededRead-only access to thirty days of web server logs. No credentials, no source code, no customer data
You receiveBot classification table, content-feeding matrix, spoofing report, citation attribution map, and a written list of the five highest-impact structural changes in priority order
PrecedentRedStar and Mykonos — both audits published in full
02 // citation architecture

The rebuild, not just the diagnosis.

from €7,500 · fixed · approx. 4–6 weeks
ForTeams with the audit findings in hand who want the structural work done and verified, not delivered as a slide deck
What happensFull AI Crawl Audit, plus restructuring of Schema.org microdata (DefinedTermSet, ScholarlyArticle, TechArticle as required), entity reconciliation against Wikidata and OpenAlex, RAG extraction optimisation across the pages that matter, and a deployment window with thirty-day verification
Access neededServer log access as above, plus either staging deploy access or working with your engineering team on handover
You receiveThe audit report, the deployed structural changes, a before/after content-feeding matrix at day 30, and an explicit statement of what the rebuild could not achieve
Free firstRun the seven-day log sample before commissioning this. If the sample returns nothing that concerns you, do not buy the paid version
03 // grounding retainer

Monthly monitoring and one surgical fix.

from €2,500 / month · rolling · cancellable any month
ForTeams who have completed an audit or rebuild and want the changes held against drift, plus a rolling fix cycle for the weakest page
What happensMonthly crawler log review, citation drift report versus previous month, entity graph integrity check, and one surgical page refresh targeting the page with the lowest Citation Share for the highest-intent query cluster
Access neededOngoing read-only server log access and a monthly fifteen-minute call to confirm priorities
You receiveMonthly PDF with bot classification, citation delta, entity graph status, plus the deployed surgical fix for that month
CancellationRolling monthly. No annual lock-in, no exit clause, no retention call. Cancel by email, effective at the next billing cycle
How pricing works

Fixed sum, agreed in writing before the engagement starts, invoiced on delivery of the report. No hourly rates, no change orders unless you change the scope. If during scoping it becomes clear the engagement will not tell you anything useful — for example, if you cannot access raw logs — I say so, and we stop there. That conversation is free.

// what i don't do

The boundary is part of the offer

An audit that promises everything cannot be checked on any of it. These are outside scope, and if one of them is what you need, the correct supplier is named — I will say so rather than take the work.

// contact

Send seven days of logs. Get a free classification report.

Before any paid engagement, there is a free version of the diagnostic. Send seven consecutive days of raw server logs — read-only, anonymised at the IP level is fine — and I return a written bot classification report: which crawlers hit which pages, three classes separated, spoofing excluded, and one sentence per page on whether it is doing anything in the retrieval layer or is simply sitting there.

No call required. No NDA required to receive the report. If it turns out you have a real problem, the report will name it in a way that lets you decide on your own terms whether to commission the paid version. If it turns out you do not, you have lost nothing except a log export.

Sample — bot classification · 7-day window · one domain
PerplexityBot 412 hits 18 pages 3 classes OAI-SearchBot 287 hits 14 pages 2 classes Claude-SearchBot 89 hits 9 pages 1 class ChatGPT-User 34 hits 6 pages 1 class Googlebot 1,204 hits 31 pages organic ── excluded ───────────────────────────── Spoofed "GPTBot" 634 hits 41 pages attack traffic Spoofed "ClaudeBot" 198 hits 22 pages attack traffic ───────────────────────────────────────── Live/RAG citations earned this window: 4 Pages fetched but not cited: 7 Pages cited but not recently fetched: 3

Request the free report WhatsApp: +48 733 123 915 ORCID

Łukasz Diener · Kraków · independent, no institutional or vendor affiliation · six published structural audits · ORCID 0009-0006-6103-8514 · OpenAlex A5136472501

"I do not ask you to believe me. I ask you to test it. Open the log. Run the classification. If your numbers differ from mine, I will retract."

— Łukasz Diener, AI Crawl Forensics, 2026
// about this page
Independent GEO consulting by Łukasz Diener

This service is not offered by an agency. It is offered by an independent researcher whose published work — six structural audits on Zenodo with permanent DOIs — uses the same method applied here to commercial sites. The RedStar and Mykonos case studies are drawn from live server logs the author holds and can show. No promises, no dashboards, no retainer traps. Just the diagnostic and the fix.

ORCID 0009-0006-6103-8514 OpenAlex LinkedIn Google Scholar All six audits Full profile