GeoScanby SEO7.es

The first 30% rule: where AI actually reads your page

Updated:

Quick answer: 44.2% of all AI citations are pulled from the first third of a page's text. Not from the best paragraph, not from the conclusion you polished for an hour. From the top. If there is no self-contained answer of 40 to 80 words inside that first third, the model has nothing to lift, and it moves on to a page that does.

That single number changed how we score content in our own auditing engine, and it explains a pattern most site owners find baffling: a genuinely good article that never gets quoted.

Why does the top of the page matter more than the rest?

Retrieval systems do not read your page. They chop it.

Before a language model can use anything you wrote, the page goes through chunking: the text is split into fragments, each fragment is turned into a vector, and retrieval happens on fragments. The model never sees your article as an article. It sees a pile of pieces, and it picks the pieces that look like an answer.

This has three consequences that most content advice still ignores.

A wall of text is a broken chunk. When a paragraph runs long, the split lands in the middle of a thought. The fragment that comes out is half an argument with no subject. Retrieval either paraphrases it, which loses your wording and your link, or discards it.

Self-contained beats well-written. A paragraph that makes sense on its own gets lifted whole. A paragraph that depends on the three paragraphs above it does not, because those three paragraphs are in different chunks.

Position is a proxy for confidence. Systems that must choose between fragments lean toward the early ones, because that is where documents usually state what they are about. The 44.2% figure is the visible result of that bias.

Chunk
A piece of text a retrieval system works with, usually one paragraph or a few in a row.
First third
The top portion of a page, the part 44.2% of all AI citations are pulled from.

See where your own page stands

172 checks in 8 groups, ten seconds, no sign-up.

Check your page

What does "first 30%" mean in practice?

The rule is not about the first 30% of your word count in some abstract sense. It is about what a machine finds when it starts reading.

How a page is split into fragments and where citations come from
Retrieval splits the page before ranking it. 44.2% of citations come from the first third.

Concretely, the marker is a direct answer of 40 to 80 words, placed either inside the first third of the plain text or in the first paragraph under your <h1> or first <h2>. Journalists call this BLUF, bottom line up front. It is the opposite of the structure most marketing content uses, where the answer arrives after three paragraphs of context-setting.

Here is the test. Take your page, delete everything except the first third, and hand it to someone who has never seen your site. Can they state what the page answers? If not, the model cannot either.

The thing I see most often is a thought smeared across the page. The answer is there, but it assembles from four paragraphs spread top to bottom. A human will read on and put it together. A model takes the first third and leaves.

Evgenii Slepinin, founder of SEO7, systems architect

Which paragraph length survives chunking?

Alongside position, length decides whether a fragment stays usable.

Paragraphs of roughly 40 to 120 words survive chunking intact. Below 40 words a paragraph rarely carries a complete thought, so it gets merged with its neighbours and loses its edges. Above 120 the split starts landing mid-argument.

This is the least glamorous advice in content optimisation and one of the most reliable. You are not writing for readability scores. You are writing units that must stand alone after being cut out of context by a machine.

Our engine measures this as a ratio: what share of your paragraphs fall in the usable band. A page where most paragraphs are two or three sentences scores well. A page built from long, elegant blocks scores badly, no matter how good the prose is.

What does this not fix?

Here is where most articles on this subject overpromise, so let me be direct.

Structuring your first third correctly does not get you cited. It makes you citable. Those are different states, and the gap between them is where most sites live.

We measured this on our own site. Five pages of seo7.es, scanned with our v1.2 engine on 27 August 2026:

Page Citability Technical Content
Home 63 82% 63%
Services 59 84% 55%
Pricing 66 85% 66%
Blog 57 85% 53%
Contacts 59 85% 55%

The technical column is nearly flat at 82 to 85%. The content column swings from 53 to 66, and it drives almost the entire spread. The answer-first check failed on the home page, and it sits at the top of the engine's "fix this first" list, worth four points on its own.

Four points. Not forty.

The same day we asked four language models (Claude, Gemini, Llama 70B and DeepSeek) whether they would name our domain when a client asks for an agency in our field. Twenty-four answers. Zero mentions.

Our pages are structured well enough to be extracted. The models do not know we exist. Those are two separate problems, and fixing the first does not touch the second.

Why does the advice still hold?

If structure alone does not get you cited, why bother?

Because it is the half you control completely.

Brand memory, whether a model knows your name at all, moves slowly. It comes from mentions on sites you do not own, from reviews, from being discussed. You can influence it, but not this week.

Extractability moves in an afternoon. Rewrite your opening so it answers the question in fifty words. Break the walls of text into units. Put the summary above the fold instead of below it. These are edits, not campaigns.

And the order matters. A model that decides to cite you and then finds nothing liftable in your first third will quote someone else. Being structurally ready is what makes brand recognition pay off when it finally arrives.

Which three checks are worth running today?

Does your first third contain a standalone answer? Forty to eighty words, no pronouns pointing at earlier paragraphs, no "as we mentioned above". If a reader landing cold on that paragraph gets a complete answer, you pass.

Do your paragraphs survive being cut out? Take three at random from the middle of the page. Read each alone. If two of them are meaningless without their neighbours, your chunking is working against you.

Is your summary above or below the fold? Many sites put a TL;DR at the end, as a recap. For retrieval that is the worst possible position: it is in the last chunk, after everything that would have made it findable.

What does a chunk actually look like?

It helps to see the difference on a real pair of paragraphs.

Written for a human reader:

We have been working in this field for a number of years now, and over that time we have seen the landscape shift considerably. What used to work no longer does, and the reasons for that are worth unpacking, because they go to the heart of how these systems have evolved and what they now reward.

Cut this out of context and it says nothing. No subject, no claim, no number. A retrieval system that lifts this fragment has lifted air.

Written to survive chunking:

AI systems pull 44.2% of their citations from the first third of a page. The reason is mechanical: retrieval splits pages into fragments before ranking them, and early fragments usually state what a document is about.

Same register, same tone, but this one stands alone. It carries a number, a claim and the mechanism behind it. Lifted into an answer, it still makes sense, and it still needs attribution, which is where your link comes from.

The difference is not writing skill. The first version is arguably better prose. The difference is whether the paragraph can be removed from your page and still work.

Frequently asked questions

Does this apply to long-form content too? More so. A three-thousand-word guide has more chunks, which means more chances for a fragment to be extracted, and more chances for a broken one to be discarded. The first-third rule holds regardless of length, because it is about where retrieval looks first, not about how much you wrote.

What if my page answers several questions? Then it has several first thirds in miniature. Each <h2> section should open with its own standalone answer. Retrieval works on fragments, and a section is closer to a fragment than a page is.

Should I delete my introduction? No. Move the answer above it. Context and background still matter for the human who stayed. They just should not stand between the reader and the point.

Do short paragraphs hurt readability for people? Two to three sentences is the length most style guides already recommend for web text. This is one of the rare cases where optimising for machines and optimising for humans point the same way.

Is a TL;DR block enough on its own? It helps, and our engine scores it separately, but it is not a substitute. A summary tells the model what the page claims. A standalone answer inside the body gives the model something to quote.

How do I know whether it worked? You will not see it in rankings, because this is not a ranking factor. You see it by asking the models directly whether they name you, and by watching whether your wording shows up in answers. We built our jury of four models for exactly that.

What did we change in our own scoring?

The 44.2% figure is one of the few numbers in this field that comes from measurement rather than from someone's blog post, which is why it carries real weight in our model: 90 points out of 3405, among the heaviest content signals we track.

But we also learned to separate what it predicts from what it does not. In our current engine, page structure and brand memory are two independent numbers, not one blended score. A page can be perfectly extractable and still invisible, and a single number hides exactly that case.

If you take one thing from this: the first third of your page is not an introduction. It is the part that gets read by machines, and everything below it is the part that gets read by humans who already decided to stay.


Sources: research digest compiled from 165 sources on AI search (August 2026); GEO-bench measurements; field scans of seo7.es with the GeoScan v1.2 engine, 27 August 2026. Details in sources/research-notes.md.