GeoScanby SEO7.es

Our site scored 96 out of 100 and got zero AI citations

Updated:

Quick answer: we built an AI-readiness auditor, ran it on our own site, and got 96 out of 100. Then we asked four language models whether they would name us when someone asks for an agency in our field. Twenty-four answers. Zero mentions. The score was not wrong about what it measured. It was wrong about what it claimed to predict, and rebuilding it around that gap took most of a month.

What was the 96 made of?

The first-generation engine checked the things everyone checks, and our site passed almost all of them.

Crawlers allowed in robots.txt, including the AI-specific user agents. Server responses fast. Pages rendered without requiring JavaScript, because we prerender. Schema markup present and valid. Canonical tags correct. Sitemap complete and submitted. Headings in a sensible hierarchy. Mobile layout clean. No broken internal links.

Every one of those is a real property of the site, correctly measured. The engine was not malfunctioning. It was reporting, accurately, that our site is technically competent.

The number then presented that as readiness to be cited by AI, which is a different claim entirely, and one the checks never tested.

Jury of models
Running the same set of questions through several language models to see whether any of them names the site.
Blocker
A condition that zeroes the whole score regardless of everything else, such as noindex or a robots.txt ban.

See where your own page stands

172 checks in 8 groups, ten seconds, no sign-up.

Check your page

Which test broke it?

We asked four models directly: Claude, Gemini, Llama 70B and DeepSeek. Six questions each, phrased the way a potential client would phrase them. Which agencies would you recommend for this kind of work in this market. Who does this well. Name some providers.

Twenty-four answers came back. Our domain appeared in none of them.

The models named competitors. They named large international agencies. Some of them named companies that, by our own engine's standards, have considerably worse technical setups than ours.

That last detail is the one that settled the argument internally. It was not that everyone scores high and we happened to lose. Sites with slower responses, messier markup and no prerendering were being recommended, and we were not.

AI is a new system for search and for being mentioned, and right now it is walking the same road organic search walked twenty years ago. We are finding out by experiment what works and what does not. Zero out of twenty four answers at 96 points was one of those experiments: the score was measuring the old system.

Evgenii Slepinin, founder of SEO7, systems architect

What was actually missing?

Not a technical property. A prior.

The models had never encountered our brand in a context that would make it available as an answer. There is no volume of mentions on sites we do not own, no accumulation of reviews and listings and discussion, nothing that would put the name into the pool a model draws from when asked to name providers.

Our pages were extractable. Nothing was extracting them, because nothing had a reason to reach for them in the first place.

Those are two separate conditions and we had been measuring one while reporting on both.

What did the rebuild change?

Three changes came out of this, and they are the substance of the current engine.

Two numbers instead of one. Citability, which measures whether a model can use your page, and memorability, which measures whether it has any prior about your brand. They are reported separately. Blending them produces a number that cannot distinguish "perfectly extractable and unknown" from "known and half-broken", which are opposite problems.

Multiplication instead of addition. The score is now content × (0.6 + 0.4 × access). Under addition, a site could accumulate points from technical checks while having nothing worth citing on the page. Under multiplication, weak content caps the whole score no matter how clean the infrastructure is. Nine specific conditions zero the score outright, because a page that cannot be fetched has no partial credit to award.

Direct testing instead of inference alone. The jury of four models is now part of the audit rather than a thing we did once to check ourselves. The engine no longer only predicts whether a model would cite you. In full mode, it asks.

What does the same site score now?

Rescanned on 27 August 2026 with the current engine:

Metric v1.1 v1.2
Score 96 63
Memorability not measured 47
Jury mentions not tested 0 of 24
Technical half not separated 82%
Content half not separated 63%

Sixty-three is not a better number. It is a more honest one, and the difference between 96 and 63 is almost entirely content: no attributed quotes anywhere on the site, no FAQ sections, statistics only on the pricing page where prices forced them, and the answer-first check failing on the home page.

The 82% technical half is the old 96 in its correct place: a component, not a conclusion.

What are the uncomfortable admissions?

Several things about this are worth stating rather than smoothing over.

We shipped the 96 to clients before we caught it. The engine was in production. Reports went out with high numbers attached. Those reports accurately described technical readiness and implied something broader.

We are still at 63 and zero. Publishing this does not fix our own memorability. The jury still returns nothing. We know what the work is and it takes months, and we are at the beginning of it.

Only 10 of 75 variables have thresholds from published measurements. The rest are informed judgements. The current model is a better-grounded hypothesis than the last one, not a validated instrument, and calling it validated would be repeating the same mistake in a new place.

The rebuild was triggered by one observation, not a study. A 96 that predicted a zero. That is the crudest possible evidence and it was sufficient, because a score that inverts its own prediction has failed regardless of sample size.

Why are high technical scores systematically misleading?

There is a structural reason this happens to good sites specifically, and it is worth understanding beyond our particular case.

Technical checks are easy to automate and cheap to run. They produce clear pass or fail results. Every tool in the category measures them, because measuring them is tractable.

Brand recognition is hard to automate, expensive to measure, and produces answers that are uncomfortable to deliver. Most tools do not measure it, and a tool that does not measure something cannot report it as missing.

So the entire category converges on scoring what is measurable and presenting it as what the client asked about. A site that is technically well built gets a high number from every tool in the market, and none of them are testing the thing that decides whether it gets recommended.

The site that suffers most from this is a competent one, because a competent site maximises the half that is measured and gets a report saying it is nearly finished.

What do you do with a high score?

If a tool has told you your site is 90-something and nothing has changed in how AI systems talk about you, the practical next steps are short.

Ask the models yourself. No tool required. Open Claude, ChatGPT, Gemini and ask what you would ask if you were a customer. Do not name your company in the question. Record whether the name comes back. This takes ten minutes and it is the single most informative thing you can do.

Separate the two problems. If the models do not know you, your website is not the bottleneck and more technical work will not move it. If the models mention you but misdescribe you, that is a content problem and it is fixable on your own pages.

Check the content half specifically. Attributed quotes, concrete figures, FAQ sections, a standalone answer near the top of each page. These are where a technically strong site is usually weakest, because they require writing rather than configuration.

Treat brand work as a separate track. Mentions on domains you do not own, listings, reviews, being discussed. Slow, unglamorous, and the only thing that moves the number that matters.

Frequently asked questions

Was the old engine badly built? It was correctly built for the wrong claim. Every individual check worked. The failure was in what the aggregate number was presented as meaning.

Does a low score mean the site is bad? No. Our own 63 comes from a technically strong site with thin content signals. The score describes citability, not quality of work or of business.

Why four models rather than one? One model's silence could be a quirk of its training data or retrieval. Four independent silences is a fact about the brand rather than about any model.

How long does memorability take to move? Months. It depends on activity outside your own domain, and there is no version of it that completes in a sprint.

Should I stop doing technical work? No. Technical readiness is the precondition. A model that decides to cite you and finds your page unusable will quote someone else. It is necessary and it is not sufficient.

Would you publish a client's low score? We publish our own, which is the only version of that question we get to answer unilaterally. The reason for publishing it is that a vendor reporting 96 on itself while invisible would be making the same error in public.

The short version

A high AI-readiness score usually means your site is technically competent. It does not mean models will name you, because those are separate conditions and most tools measure only the first.

The test that settles it takes ten minutes and costs nothing: ask the models what a customer would ask, without naming yourself, and see whether you come back. We did, and we did not. Everything we rebuilt afterwards came from that.


Sources: GeoScan v1.1 and v1.2 scans of seo7.es, 26 and 27 August 2026; four-model jury tally (Claude, Gemini, Llama 70B, DeepSeek), 24 answers, same dates; GEO-bench measurements (Princeton University); project research digest compiled from 165 sources on AI search, August 2026. Details in sources/research-notes.md.