GeoScanby SEO7.es

Version history: v1.0 to v1.2, including the score that fell from 96 to 63

Updated:

Quick answer: three versions in five weeks. The first scored pages and called it AI readiness. The second asked models directly and found the gap. The third was rebuilt around it, and the same site that scored 96 now scores 63. The lower number is the honest one, and this page explains why we published the drop instead of quietly rescaling.

What changed in v1.2, 26 August 2026?

What changed.

There are two answers now: whether the page is ready to be cited, and whether AI knows the brand at all. One number could never explain why a technically flawless site goes uncited.

Blockers arrived. A site excluded from the index, disallowed in robots.txt, or serving a captcha to bots drops to zero rather than losing a few points, because such a page cannot be cited at all.

Scoring became proportional. Alt text on 80% of images is 0.8, not a failure.

Every check now explains how it should be and what was actually found. Grouped checks name each blocked bot and each missing tag instead of reporting a ratio.

You choose the depth: a quick scan of the page itself, or a full one that adds load speed, brand mentions and reviews on third-party platforms.

The characteristic list was folded into 75 weighted variables, because an obscure crawler used to weigh as much as an outright ban on citation.

What it cost us. Our own home page went from 96 to 63. Nothing about the site changed between those two scans.

Engine version
A set of weights and thresholds fixed on a date. Old reports are not recalculated under a newer version.
Score regression
A drop in the number caused by a change in how it is measured, not by the site getting worse.

See where your own page stands

172 checks in 8 groups, ten seconds, no sign-up.

Check your page

Why did the same site lose 33 points?

The old score was accurate about what it measured and wrong about what it claimed.

It measured technical readiness: crawlers allowed, fast responses, valid schema, prerendered HTML, correct canonicals. Our site passes all of that, which is why it scored 96.

Then we asked four models, six customer-style questions each, whether they would name us. Twenty-four answers, zero mentions. Some of the companies they did name have messier markup and slower responses than ours.

A score of 96 that predicts zero is not slightly off. It is measuring one thing and reporting another.

Metric v1.1 v1.2
Score 96 63
Technical half not separated 82%
Content half not separated 63%
Memorability not measured 47
Models naming the brand not tested 0 of 24

The 82% is the old 96 in its correct place: a component, not a conclusion. The full account is in Our site scored 96 and got zero citations.

What did we change in the report itself, 29 August 2026?

Five checks used to appear as a single line each, without their contents. Trust signals said "4 of 12" without naming which eight were missing. Headings, the author chain, compression and the speed metrics behaved the same way.

They now list every item with its own verdict, which raised the number of visible checks from 148 to 172 without adding a single new check. The engine was already measuring all of it and showing only the summary.

The number in our materials moved from 305 to 172 at the same time. Both are true for their own version: v1.0 really did run 305 checks across seven categories. The landing page simply kept the old figure after the engine was rebuilt, and a tool that sells accuracy cannot leave that standing.

What changed in v1.1, 5 August 2026?

The Reality block began asking the model six different customer questions instead of one, and showing who it names instead of you.

The matrix became 3×2: "not named, sometimes, often" on one axis and page readiness on the other.

The report was rewritten as tables. Every row got a description in the visitor's language and a result as an icon.

The result became visible immediately, with no email required. Email is only needed for the PDF version.

What shipped in v1.0, 22 July 2026?

305 checks across seven categories: AI bot access, structured data, content, citability, meta tags, technical base and trust.

A question to the language model: does it name this site when answering a customer.

Three languages, a PDF report by email, no signup.

What did we not change, and why?

The engine still lives in two places. The site runs it inside an n8n pipeline, the API runs it in the backend. Rather than trusting them to stay identical, the pipeline nodes are generated from the same source files, and a test executes both and compares every observation. The previous version spent three weeks with a signature that disagreed with its own code.

Old scans keep their old numbers. History is not rescaled to the current model. A report issued in July says 305 checks because that is what ran in July.

The v1.1 pipeline is still alive as a rollback. It will be retired once the new one has run without incident for a week.

What is coming, honestly labelled?

Validation. The weights are informed judgements: ten thresholds of seventy-five come from published measurements. Until we accumulate scans and correlate citability with what models actually say, the model stays a hypothesis. This is the largest open item and it has no date.

More sources for memorability. Five live sources today, four of them from one search engine. Maps, catalogues, referring domains and structured knowledge bases are the obvious gaps.

Four checks scored at zero. llms.txt, a markdown twin, content negotiation and an MCP endpoint. They are watched, reported separately and weighted zero until a provider commits to reading them. See We gave llms.txt a weight of zero.

Frequently asked questions

Will my old report still match a new scan? No, and it should not. Rescan after a version change: the numbers are not comparable across engines, which is why every report carries its engine version.

Why publish a score drop instead of rescaling quietly? Because clients holding a 96 would otherwise discover the gap themselves. A vendor that reports 96 on its own site while four models return nothing is making the exact error it sells a tool against.

Is a lower score worse for me? It is more accurate. A v1.2 score of 63 and a v1.1 score of 96 describe the same site. Only one of them told you that no model knows your brand.

How do I know the site and the API agree? They are generated from the same source and compared by an automated check on every build. The report carries a hash of the check set, so two reports with the same hash ran the same engine.

What triggers a new version? Evidence that the current model is wrong, not a schedule. v1.2 was triggered by one observation: a 96 that predicted a zero.

Where do the measurements come from? What GeoScan measures lists every figure with its source and marks which thresholds are published research and which are our judgement.