GeoScanby SEO7.es

The denominator problem: nobody can measure AI visibility and most tools pretend otherwise

Updated:

Quick answer: to state a share of AI visibility you need to know how often you were cited and how often you could have been. The second number does not exist. Nobody knows how many relevant conversations happen inside AI assistants, what was asked, or what was answered. Every percentage of "AI share of voice" you have seen has an invented denominator underneath it.

What did we lose when search became conversation?

Classical search leaks its own measurements. Query volume is estimated by tools with access to clickstream data. Impressions and clicks arrive in Search Console. Position is checkable by anyone who runs the query. The denominator is imperfect but it exists, and the whole discipline of SEO reporting is built on it.

AI assistants leak nothing.

Conversations are private. Nobody publishes how many people asked about your category this month, or how those questions were phrased, or what proportion of answers named anyone at all. There is no equivalent of Search Console, no aggregate query volume, and no impression count.

So the question "what share of AI answers mention us" has a numerator you can sample and a denominator that does not exist. Reporting a percentage requires manufacturing the missing half.

Share of voice
The share of brand mentions among all mentions on a topic.
Denominator
The total number of chances to be cited, the number that does not exist for AI search.

See where your own page stands

172 checks in 8 groups, ten seconds, no sign-up.

Check your page

How does the missing number get manufactured?

Three methods circulate, and each produces a number that looks precise and is not.

A fixed prompt set. The tool asks a hundred prompts it chose, counts how often you appear, and reports the ratio. The denominator here is not "relevant conversations" but "prompts we picked", and the number moves entirely with the prompt selection. Change the list and the score changes, with no change to reality.

Keyword volume as a proxy. Take classical search volumes and assume AI queries distribute similarly. They do not. Conversational questions are longer, more specific, and often have no search-keyword equivalent. This substitutes a measurable denominator for the unmeasurable one and hopes nobody checks the substitution.

Extrapolation from a sample. Run some prompts, observe a rate, project it. This is legitimate if the sample is representative, and there is no way to establish that a sample of prompts is representative of a population you cannot observe.

None of the three is fraudulent. All three produce a number whose precision is an artefact of the method rather than a property of the world.

What can actually be measured?

Being honest about the denominator does not mean measuring nothing. It means being specific about what each measurement is.

Presence on a fixed question set. Ask a defined list of questions, record whether the brand appears, repeat the same list over time. This is a real measurement. It just is not a share, it is a count on a set you defined. Its value is in the change between runs, not in the absolute number.

Agreement across models. Ask several independent models the same questions. If all of them return nothing, that is a robust finding, because the failure modes of different models are not identical. Our own tally is four models and twenty-four answers, and the zero across all four says more than any single model's zero would.

Extractability of your pages. Fully measurable, entirely on your side, no denominator required. Whether your content can be chunked, whether an answer exists in the first third, whether claims are attributed.

External mention counts. Countable, with caveats about coverage. How often the brand appears on domains you do not own is a real number, even if it is not the same as the number a model has internalised.

The common property: each of these is a measurement of something, stated as that thing rather than as a proxy for market share.

Which trap did we walk into ourselves?

A worked example of how easy the mistake is, from our own work.

We measure a memorability figure from external signals, and on our first run it came back as 28. That number went into reports as a finding.

Then we discovered a bug: search operators in our queries were breaking the source we were pulling from, so the query returned advertisements instead of organic results. After the fix, external mentions went from 0 to 13, forum mentions from 0 to 10, and memorability from 28 to 47.

Nothing about the brand changed between those two runs. Nineteen points of difference came from a query-formatting error.

The lesson we took: any measurement whose inputs come from scraping is a measurement of your scraper as much as of the world, and a number that moves nineteen points on a bug fix should be reported with that fragility attached.

This is also why our engine returns null for memorability when it has fewer than three external signals, rather than returning a low number. A low number implies "we looked and found little". A null says "we did not gather enough to have an opinion". Those are different statements and collapsing them is how invented denominators get into reports.

What does honest reporting look like?

Four properties, offered as a standard we try to hold ourselves to and do not always hit.

Say what the denominator is. If the score is over a prompt set, name the prompt set. If it is over external sources, name the sources. A percentage with an unnamed denominator is a claim about the world made from a sample of unknown relation to it.

Separate observation from inference. "Zero of twenty-four answers named the brand" is an observation. "You have low AI visibility" is an inference. Both are useful; only one is a fact.

Report the fragility. If a number would move on a bug fix, a source outage or a prompt rewording, say so. Precision that survives no perturbation is decoration.

Say when you do not know. The hardest one commercially, because a report that admits gaps is a worse sales object than one that does not. It is also the only kind that stays true when the client checks.

Giving a client a pretty percentage means signing your name under it. Preparing a site to be cited is something we can do. Guaranteeing it is something nobody can do: the model is a black box, and all we hold are testable hypotheses. I would rather hand over a hypothesis with reasoning than a number without it.

Evgenii Slepinin, founder of SEO7, systems architect

Why does this matter more than it sounds?

The denominator problem is not a technicality about reporting hygiene. It shapes what the entire field optimises for.

When a number can be manufactured, it gets optimised. Tools compete on producing numbers that move, because a client seeing movement renews. The easiest numbers to move are the ones with invented denominators, because the denominator can be adjusted.

This is how a discipline drifts into measuring its own instruments. The technical checks get refined, the scores get more decimal places, the reports get longer, and none of it touches whether anyone is actually being recommended by an AI assistant.

The corrective is boring and available to anyone: ask the models what a customer would ask, without naming yourself, and write down what comes back. It costs nothing, it has no denominator problem because it makes no share claim, and it is the only measurement in this field that cannot be gamed by the person being measured.

What can we still not tell you?

Stated plainly, because an article about honest measurement should demonstrate it.

We cannot tell you what percentage of AI answers in your category mention you. Nobody can.

We cannot tell you how many people asked an assistant about your category this month.

We cannot tell you whether a specific change to your site caused a change in citations, because we cannot isolate the variable in a system we cannot observe.

We can tell you whether four models name you on a defined set of questions, whether your pages are extractable, and how many external mentions we could find. That is a smaller claim than the category usually makes and it is the size of claim the available evidence supports.

Of the 75 variables in our own model, ten have thresholds from published measurements. The other sixty-five are informed judgements. We say so in the reports, and it makes them harder to sell.

Frequently asked questions

Is AI share of voice completely meaningless? Not meaningless, but it is a measurement of a prompt set rather than of a market. Useful for tracking change over time against a stable list. Not useful as an absolute figure or as a comparison across tools that use different lists.

Can I build my own prompt set? Yes, and it is the best version of this. Use the questions your customers actually ask, keep the list fixed, run it monthly. Your list is more relevant to you than any vendor's generic one.

How many prompts is enough? Enough to cover the distinct things customers ask, which for most businesses is somewhere between ten and thirty. More prompts do not fix the denominator problem, they just make the sample bigger on a population you still cannot see.

Will the denominator ever exist? Only if providers publish aggregate data, which they have no obligation and little incentive to do. Plan on it not existing.

Does this mean AI visibility work is unmeasurable? No. It means the outcome is measurable as presence on defined questions, and the inputs are measurable directly. What is not measurable is market share.

Why publish this if you sell audits? Because the alternative is competing on invented precision, and that competition is won by whoever is least careful. Also because clients eventually check.

The short version

Every AI visibility percentage rests on a denominator nobody has. Conversations with assistants are private, query volumes are unpublished, and no amount of tooling recovers a number that was never emitted.

What remains is smaller and true: ask the models a fixed set of customer questions, record whether your name comes back, measure whether your pages can be used, and count mentions on domains you do not own. Then say those are what you measured, rather than dressing them as a share of a market you cannot see.


Sources: four-model jury methodology and tally, seo7.es, 27 August 2026; memorability measurement bug and correction (28 → 47), same date; GeoScan v1.2 variable definitions; project research digest compiled from 165 sources on AI search, August

  1. Details in sources/research-notes.md.