AI Search

How Does Google Choose AI Overview Sources?

10 August 2026 11 min read
Short answer

Google does not publish a fixed formula for choosing AI Overview sources. Its public documentation says AI Overviews use Google's core Search ranking and quality systems, while a technique called query fan-out can issue several related searches across subtopics and data sources. Google's models then identify supporting pages and generate an answer with links.

What decides which pages Google cites in AI Overviews?

Google does not publish a fixed formula for choosing AI Overview sources. Its public documentation says AI Overviews use Google's core Search ranking and quality systems, while a technique called query fan-out can issue several related searches across subtopics and data sources. Google's models then identify supporting pages and generate an answer with links.

That means the cited pages do not have to match the ordinary top ten for the original query. A source can be selected because one passage supports one part of the generated answer, even when the page targets a broader or narrower subject. This is the mechanics our AI Overview optimisation work is built around.

There are three types of evidence worth separating throughout this discussion:

Google confirms: statements in Google's official documentation.

Research observes: patterns found in a disclosed dataset.

Reasonable inference: a practical interpretation that has not been confirmed as a ranking rule.

Confusing those categories is how sensible observations become fictional "AI ranking factors".

Key takeaways

  • Google says AI Overviews rely on core Search ranking and quality systems, plus query fan-out across subtopics.
  • Cited pages do not have to match the ordinary top ten for the original query.
  • Foundations still matter: crawlability, Search policy compliance, and helpful, reliable, people-first content.
  • A useful passage answers a specific part of the question, contains support rather than assertion, and survives extraction.
  • Reported top-ten overlap with citations has varied widely across studies — from 76.1% to 38% — so treat any single figure with caution.
  • About 11% of claims in one large study were unsupported by the cited pages, mostly through omission.
  • Publishers can control eligibility, clarity, evidence and authority, but cannot reserve a citation.

AI Overviews retrieve supporting material for a generated answer

Traditional Google results normally present a ranked set of pages for the query. An AI Overview has a different job: it attempts to answer the question and help the user explore further.

Google says AI Overviews and AI Mode may use query fan-out. Instead of running only the exact search entered, the system can issue multiple related searches across subtopics and data sources. Its models identify supporting pages while constructing the response.

Consider the search "best insulation for a Victorian terrace with solid walls". To produce a useful answer, the system may need material about:

  • internal versus external solid-wall insulation;
  • condensation and moisture risk;
  • planning or listed-building constraints;
  • disruption and cost; and
  • suitability for a particular property.

One general guide may not be the strongest source for every point. The overview could cite an official source for planning, a technical specialist for moisture and a practical renovation guide for disruption.

The simplified process looks like this:

  • The user asks a question.
  • Google interprets the main intent and related information needs.
  • Query fan-out retrieves candidate results for those needs.
  • The models identify passages that can support the response.
  • Google generates the answer and displays associated links.

This is a useful mental model, not a complete description of every internal system. Google does not disclose all signals, weightings or generation decisions. It is why our AI Overview optimisation specialists treat this as a probability exercise rather than a checklist to game.

What Google publicly says about inclusion

Google's current guidance is deliberately unglamorous. The same foundational practices that support Google Search also support its generative features.

The page must meet Search's technical requirements

Content must be accessible to Google, indexable and eligible to appear in Search. Google must be able to discover the page through links or other normal discovery routes, crawl it and process the main content.

The page must follow Search policies

Spam policies still apply. Generative features do not create an exemption for scaled low-value content, cloaking, link manipulation or misleading structured data.

The content should be helpful, reliable and people-first

Google advises publishers to create content for a real audience, demonstrate first-hand expertise where relevant and give readers substantial value. Its 2026 generative-AI guidance specifically recommends non-commodity, expert-led content that goes beyond common knowledge.

Important information should be available in text

Images and video can be valuable, but essential meaning should not be trapped inside media Google or assistive technologies cannot interpret. Descriptive text also gives the system clearer material to retrieve.

Images and video can strengthen multimodal usefulness

Where a process, product or destination is visual, high-quality original media can help users and Google understand it. A labelled photograph showing an installation detail contributes more than a generic stock image.

Structured data should match the visible page

Google uses supported structured data to understand page information and enable rich-result eligibility. It does not say schema is required for AI Overviews. Markup should describe what users can actually see.

No special AI file is needed

Google says it does not use llms.txt or other special AI text files for inclusion in Search's generative features. It also warns against unnecessary AI-specific chunking and inauthentic mentions.

These statements establish the foundations. They do not reveal a scorecard for citation selection.

What makes a passage useful to a generated answer?

Google does not publish a checklist for a "citable passage". The following characteristics combine official quality guidance with reasonable editorial inference.

It answers a specific part of the question

A passage is more useful when its subject and conclusion are clear. "This depends on several factors" has little standalone value. "Most fixed-rate mortgages limit annual overpayments, but the allowance and calculation period depend on the lender and product terms" identifies the subject, general rule and limitation.

Clear does not mean simplistic. Qualifiers should remain beside the answer where removing them would make it misleading.

It contains support, not just assertion

An answer becomes stronger when it includes evidence: a current figure, calculation, source, first-hand observation or expert explanation. A claim such as "external wall insulation is always best" is both unsupported and too absolute. A useful source explains the conditions that change the answer.

It is complete enough to survive extraction

Generated systems may use a section outside the context of the full article. Name the subject explicitly. Include dates, jurisdiction and conditions when they determine whether the statement is true.

For example, "The allowance is £20,000" is incomplete. "The UK adult ISA subscription limit is £20,000 for the 2026/27 tax year" is bounded by product, market and date.

It contributes something to the source set

Different sources can play different roles. One may define the issue, another provide current data, another explain a process and another describe an exception. A page does not need to be the definitive guide to the whole subject to support one part well.

It sits on a page and site Google can evaluate as trustworthy

The quality of an isolated paragraph is not the only consideration. Google's systems can use page-level and site-wide signals. Transparent authorship, primary sources, relevant expertise, reputation and technical reliability all help a reader — and potentially Google — judge the material.

Our guide to E-E-A-T and AI Overviews explains how to make that trust verifiable. It sits alongside our wider work on what AI SEO actually involves, since the two disciplines share the same evidentiary standards.

A source-passage comparison

Imagine three pages discussing mortgage overpayments.

Page A says: "Overpaying your mortgage is a great way to save money. Contact us to learn more."

It is promotional, vague and contains no conditions.

Page B says: "Many lenders let borrowers overpay up to 10% each year."

It provides a useful general pattern but lacks a warning that lender policies and calculation methods differ.

Page C says: "Many UK fixed-rate products permit limited penalty-free overpayments, often expressed as a percentage of the outstanding balance. The allowance, measurement period and early repayment charge are product-specific, so check the current mortgage offer before paying."

Page C is more useful because it explains the pattern, variables and action. Add an example using real product terms, a current review date and a qualified reviewer, and it becomes more defensible still.

This does not prove Google will cite Page C. It shows how to create a passage that better supports an accurate answer.

Passage quality compared
PageContentUsefulness
Page APromotional, vague, no conditionsLow — no support, no specifics
Page BGeneral pattern, no caveatsMedium — useful but incomplete
Page CPattern, variables and action statedHigh — specific, supported, actionable

Why cited sources can differ from page-one results

Page-one ranking and AI Overview citation overlap, but they are not the same system output.

In July 2025, Ahrefs reported that 76.1% of AI Overview-cited pages in its study also ranked in the top ten for the same query. In March 2026, after analysing 863,000 SERPs with a revised contemporary dataset, Ahrefs reported a top-ten overlap of 38%.

A separate 2026 academic study examined 55,393 trending queries across 19 categories. It found that nearly 30% of AI Overview-cited domains did not appear among the co-displayed first-page results.

These figures should not be blended into a universal percentage. They were produced at different times using different samples and methodologies. They do, however, support two sensible conclusions:

  • Traditional ranking strength remains relevant.
  • The pages cited by an AI Overview are not limited to the visible top ten for the original query.

Query fan-out offers a clear explanation. A specialist page might rank poorly for the broad query but strongly answer a narrow related search used to support one sentence in the synthesis.

Read our full examination of whether page-one rankings are required for AI Overviews. For a broader look at surfaces beyond Google, see how to rank higher in ChatGPT.

How might a claim become associated with a source?

Google's interfaces change, and citations are not always a simple footnote attached to one sentence. The safe editorial principle is to ensure every factual claim is directly supported by the source linked near it.

When preparing an article, use a claim-evidence check:

  • Highlight each factual claim.
  • Identify the exact source passage supporting it.
  • Check the source's date, jurisdiction and authority.
  • Confirm the article does not state more than the evidence supports.
  • Place the citation close enough for readers to understand the relationship.

This matters because citation systems are not infallible. The 2026 study of 55,393 queries decomposed AI Overviews into 98,020 claims and reported that 11% were unsupported by the cited pages, most often through omission. Publishers should make their own claim-to-evidence relationships exceptionally clear, even though they cannot control how a generated interface ultimately attributes information.

Source diversity, authority and freshness

The right source type depends on the question.

  • Use current government or regulatory material for rules and allowances.
  • Use original research for a study's findings.
  • Use a qualified specialist for professional interpretation.
  • Use first-hand documentation for a process or experience.
  • Use forums cautiously to understand lived experiences and language, checking factual claims elsewhere.
  • Use video when demonstration materially improves the answer.

Freshness is also query-dependent. The current ISA allowance needs a current source. The history of double-entry bookkeeping does not need a weekly update.

Do not change a publication date merely to make an old page look fresh. Review the content, update what has materially changed, retain what remains accurate and disclose the review date.

Authority does not always mean the largest domain. For a narrow technical question, a specialist may provide better evidence than a general publisher. For a regulated fact, the official source may be more appropriate than either.

What can publishers control?

You can improve:

  • technical eligibility;
  • page discovery and internal linking;
  • the clarity and completeness of the answer;
  • original evidence and first-hand value;
  • source quality and attribution;
  • author and reviewer transparency;
  • external reputation and corroboration;
  • structured data accuracy; and
  • ongoing maintenance.

You can also use Google's snippet and preview controls where there is a genuine reason to limit how content appears. Getting the content strategy right often means revisiting how pages are planned in the first place through content strategy and optimisation.

You cannot control:

  • whether an AI Overview triggers for every search;
  • the exact wording Google generates;
  • which related searches Google issues;
  • which source supports each part on every occasion;
  • how prominently a citation appears; or
  • how stable that selection remains.

This distinction prevents overpromising. SEO can improve eligibility, relevance, quality and authority, and learning how to get in the Google AI Overviews is really about strengthening those levers consistently. It cannot reserve a citation.

Become the useful source, not merely another result

Google chooses AI Overview sources dynamically. The best response is not to reverse-engineer a fictional static formula. Make each important page eligible, specific, well supported and valuable enough to contribute something the rest of the source set lacks.

Undeniable Growth can compare your pages with the sources currently cited across priority searches and identify the clearest evidence, content and authority gaps through an AI source-gap analysis.

Sources

Frequently asked questions

Does Google always cite the highest-ranking page?+

No. Research shows overlap between high organic rankings and citations, but query fan-out can introduce pages outside the co-displayed top ten. A high ranking is helpful evidence of relevance and quality, not a guaranteed citation.

Can Google cite more than one page from the same website?+

It can, where different pages provide useful supporting information. A site should still avoid creating overlapping pages solely to occupy more citation space. Give each URL a distinct purpose and consolidate duplicates.

Does Google use whole pages or individual passages?+

Google retrieves and evaluates pages, but particular passages can support different parts of a response. Write coherent pages with clear sections; do not reduce an article to disconnected fragments designed for extraction.

How fresh must a source be?+

There is no universal age limit. Freshness matters when the answer changes with time, such as prices, laws, allowances and product features. Stable concepts can remain useful for years if they are accurate and well maintained.

Can forums and user-generated content be cited?+

Yes. Google can surface material from discussions when lived experience or opinion is relevant. User-generated content is not automatically reliable, so factual and high-stakes claims require careful corroboration.

Is there an official checklist for becoming a citable passage?+

No. Google has not published one. The characteristics that help — answering a specific point, containing support, surviving extraction and contributing something distinct — are drawn from official guidance combined with reasonable editorial inference.

Does structured data guarantee a citation?+

No. Google says structured data is not required for AI Overviews. Markup should accurately describe what is visible on the page; it supports understanding but does not buy inclusion.

Is an llms.txt file needed to be included?+

No. Google says it does not use llms.txt or other special AI text files for inclusion in Search's generative features.

Why did Ahrefs report different overlap percentages in 2025 and 2026?+

The studies used different samples, dataset sizes and methodologies at different points in time, so the figures — 76.1% and 38% — should not be blended into one universal statistic, though both indicate meaningful overlap alongside genuine divergence.

Conclusion

Google chooses AI Overview sources dynamically. The best response is not to reverse-engineer a fictional static formula. Make each important page eligible, specific, well supported and valuable enough to contribute something the rest of the source set lacks.

Undeniable Growth can compare your pages with the sources currently cited across priority searches and identify the clearest evidence, content and authority gaps — talk to us to get started. Portsmouth businesses can also work with our AI Overview optimisation team in Portsmouth directly.

Related reading

Turn AI search visibility into measurable pipeline.

A short review shows where your site is already close to being cited, and what to fix first.