GeoScanby SEO7.es

One score hides the problem: why AI readiness needs two numbers

Updated:

Quick answer: a single AI-readiness score cannot describe your situation, because two independent things decide whether a model cites you: whether your page can be extracted, and whether the model has ever heard of your brand. A site can be perfect on the first and invisible on the second. Blending them into one number hides exactly the case that matters.

Which failure forced the split?

Our own site scored 96 out of 100 on our first-generation auditing engine. Clean markup, fast responses, schema everywhere, crawlers allowed, sitemap valid.

We then asked four language models whether they would name our domain when someone asks for an agency in our field. Twenty-four answers across Claude, Gemini, Llama 70B and DeepSeek. Zero mentions.

A score of 96 that predicts zero is not a slightly inaccurate score. It is measuring something real and reporting it as if it were something else.

What 96 actually described was extractability: if a model decided to use our page, it would find the page usable. What it did not describe, and could not, was whether any model would decide to use it in the first place.

Citability
Whether a model can find, read and use the page as a source.
Memorability
Whether a model knows the brand at all, before any page is fetched.

See where your own page stands

172 checks in 8 groups, ten seconds, no sign-up.

Check your page

What are the two axes, stated plainly?

Citability. Can a model extract and reuse what is on your page? This covers access (does the crawler reach the page, does it render without JavaScript, is the response fast enough) and content (is there a standalone answer near the top, do paragraphs survive chunking, are there figures and attributed claims).

Citability and memorability as two independent axes
Two independent axes. Our own site sits bottom right: extractable and unknown.

Citability is almost entirely under your control. It changes in an afternoon.

Memorability. Does the model have any prior about your brand at all? This comes from mentions on sites you do not own, from being discussed, reviewed, listed, cited by others. It is measured from outside your domain, and you cannot edit it directly.

Memorability moves in months, through activity that has nothing to do with your HTML.

The two are close to independent. High citability with zero memorability is our own situation and a very common one. Low citability with high memorability describes a well-known brand whose site blocks crawlers, which is also common and much easier to fix.

Why does blending them produce a useless number?

Take a weighted average of the two and you get a number that behaves badly at both ends.

A site with citability 90 and memorability 10 averages to 50. A site with citability 50 and memorability 50 also averages to 50. These are completely different situations requiring completely different work, and the blended score cannot tell them apart.

Worse, the average is optimistic for the first site and pessimistic for the second. The first site is one good year of external activity from being cited constantly. The second has two half-finished problems.

The single number also creates a bad incentive. Citability is the half you can move this week, so any blended score can be raised by working only on the half that does not solve your actual problem. You watch the number climb and nothing changes in the answers.

In medical equipment a device is not cleared for work until it has been verified against a reference standard. AI has no reference values. But a site with real expertise can itself become the reference: models have nothing to check themselves against except sources worth trusting.

Evgenii Slepinin, founder of SEO7, systems architect

What do we do when memorability cannot be measured?

An honest second number requires external evidence, and external evidence is not always available.

Our engine computes memorability only when it can gather at least three independent external signals: mentions across the web, presence on aggregator and review platforms, discussion on forums, structured data connecting the site to external profiles, and the direct test of asking models whether they know the brand.

Below three signals, the engine returns null for memorability rather than a low number.

This distinction matters more than it sounds. A memorability score of 12 means "we looked and found almost nothing". A null means "we did not look hard enough to say". Reporting the second as the first would be inventing a measurement, and a tool that invents measurements is worse than a tool that admits gaps.

What do the two numbers look like on a real site?

Our own domain, scanned 27 August 2026 in full mode:

Metric Value
Citability 63
Technical half 82%
Content half 63%
Memorability 47
Models naming the brand unprompted 0 of 24 answers

The gap between the two numbers is the diagnosis. Technical work is largely done at 82%. Content is the weaker half at 63%. And memorability at 47, with zero unprompted mentions from any of the four models, says the brand simply is not in the models' priors.

If we had reported a single blended figure of roughly 55, none of that would be visible. The number would look mediocre across the board and suggest a general tidying-up. The two numbers instead point at one specific piece of work: get mentioned somewhere that is not our own site.

How does the jury work as a reality check?

The most useful part of the second number turned out not to be the score itself, but the direct test behind it.

We ask four models the same set of questions a potential client would ask, and record whether the brand appears in the answer. Four models rather than one, because a single model's silence could be a quirk of that model. Four independent silences is a fact about the brand.

A failed API call is not counted as "did not cite". If a model times out or errors, it drops out of the tally rather than voting against you. Otherwise an infrastructure hiccup would masquerade as a finding.

The tally is blunt and it is the least gameable measurement in the whole system. You cannot optimise your way to being known. Either the models produce your name or they do not.

What does this mean if you are buying an AI audit?

A practical filter, offered without much diplomacy: if a tool reports one number and that number is high while nothing changes in actual AI answers, the tool is measuring extractability and calling it visibility.

Questions worth asking of any AI-readiness report:

Does it separate what I control from what I do not? Technical and content fixes are yours. Brand recognition is not, and a report that mixes them will overstate how much of your problem is solvable by editing your site.

Does it test the models directly? A score derived entirely from your own HTML has never asked a model anything. It is a prediction, not an observation.

Does it say when it does not know? A tool that always produces a confident number is not measuring anything that can fail to be measurable.

Does the score move for reasons you can explain? If two scans a week apart differ by fifteen points with no changes to the site, the number is noise wearing a decimal point.

What is the uncomfortable part for us?

Splitting the score made our own product harder to sell.

One number is a great sales object. It goes up, the client sees progress, the invoice makes sense. Two numbers, one of which explicitly says "this part will take months and is not primarily about your website", is a worse pitch and a more honest report.

We also have to state that our own site scores 63 and 47, which is not a flattering position for a company selling AI visibility work. The alternative was to keep reporting 96 and let clients discover the gap themselves.

Of the 75 variables in our current model, only ten carry thresholds taken directly from published measurements. The rest are informed judgements. The split into two axes is not one of the measured findings either. It came from watching a 96 predict a zero, which is the crudest kind of evidence there is and also the hardest to argue with.

Frequently asked questions

Can I raise memorability by editing my site? Only marginally. Structured data linking your site to external profiles helps a model connect what it already knows to you. But if there is nothing external to connect to, the markup connects to nothing.

How long does memorability take to move? Months, and it depends on activity outside your domain: publishing where others read, being listed, being reviewed, being discussed. There is no version of this that happens in a sprint.

Should I fix citability first or memorability first? Citability first, because it is fast and because it is the precondition. A model that decides to cite you and finds nothing extractable will quote someone else. Structural readiness is what makes recognition pay off when it arrives.

Is a citability score of 63 bad? It is middling and it is honest. The technical half is largely done; the content half is where the work sits. A score that high with a failing content half is a common shape for sites built by developers rather than writers.

Why four models and not one? One model's silence is anecdote. Four independent silences is a pattern. Models differ in training data and retrieval behaviour, so agreement across them carries more weight than any single answer.

What happens if the models start citing me? Memorability rises and, more usefully, the jury tally changes from zero to a number. That transition is the only measurement in this field that unambiguously means progress.

The short version

Two things decide whether a model names you: whether it can use your page, and whether it knows you exist. They move at different speeds, they respond to different work, and averaging them produces a number that describes neither.

The first is engineering and takes an afternoon. The second is reputation and takes a year. A report that does not tell you which one is your problem has not told you anything.


Sources: field scans of seo7.es with the GeoScan v1.2 engine, 27 August 2026; four-model jury tally, same date; GEO-bench measurements (Princeton University); project research digest compiled from 165 sources on AI search, August 2026. Details in sources/research-notes.md.