How to analyze sources used by AI Overviews

This article gives you the specific analysis workflow: how to test, what to log, how to compare cited vs non-cited pages, and how to translate findings into concrete changes.

Why analysing AI Overview sources is worth doing

When Google’s AI Overview appears for a query, it synthesises an answer from multiple web sources. Some of those sources are cited by name. The selection is not random and it does not perfectly mirror organic ranking order. Authoritas published analysis showing that pages ranked in positions four through ten are regularly cited when their content is better structured for extraction than the top-ranking pages.

The signals that earn citation are not identical to the signals that earn ranking. Citation selection rewards content structure, schema presence, and answer-block organisation in ways that traditional ranking optimisation does not require.

Source analysis closes that gap. By examining which pages get cited for a given query and what they have in common, you build a template for what your own content needs to look like to compete for citation. Rank tracking tells you if you are in the game. Source analysis tells you why you are winning or losing the citation layer.

Step 1: Manual testing methodology

Manual testing is the most reliable source analysis method currently available. AI citation monitoring tools exist, but they are either expensive, narrow in scope, or inconsistent in coverage. For most practitioners, manual testing is the starting point and the ongoing baseline.

Query selection

Choose ten to twenty queries directly relevant to your business or client. Include a mix of core service category queries (“family dentist Tucson,” “accounting firm small business Austin”), informational queries in your topic area (“how often should I see a dentist,” “what is cash basis accounting”), and comparison or evaluation queries (“best accounting software for small business,” “dentist vs orthodontist for braces”).

Avoid navigational queries (searches for a specific brand or URL) because these rarely trigger AI Overviews. Focus on informational and research queries where AI Overview prevalence is highest.

Which engines to test

Run each query in at least three systems: Google Search (look for the AI Overview at the top of the SERP), Perplexity (which shows cited sources explicitly in a sidebar), and ChatGPT using browse mode or the search-enabled version for real-time citation.

Testing across multiple engines matters because each system draws on different retrieval mechanisms. A page cited in Google’s AI Overviews may not appear in Perplexity and vice versa. Patterns that appear across multiple engines are stronger citation signals than those that appear in only one.

The logging format

Use a spreadsheet with these columns for every query you test:

Date | Engine | Query | AI Overview present (Y/N) | Cited URLs | Your URL cited (Y/N) | Competitors cited | Structural observations

The structural observations column is the most important. Do not leave it blank. Force yourself to record at least one observation about each cited page: whether it leads with a direct answer, whether it has visible FAQ content, whether the cited section is in the opening paragraphs or deeper in the page.

Run the full query set once a month. After three months, you have a trend log that shows which queries are gaining or losing citation presence and which competitors are appearing more or less consistently.

Step 2: Auditing cited sources

Once you have identified which pages are being cited for your target queries, open each one and audit it across five dimensions.

Answer-block position

Where does the relevant answer appear on the page? In most cited pages, the answer to the query question appears within the first two paragraphs or in a clearly marked section near the top. Pages where the answer is buried deeper, or only accessible by reading the full article, are less commonly cited.

Note whether the first paragraph of the cited page could stand alone as a complete, useful answer to the query. If it could, the page has strong extraction structure. If it could not, look for where on the page the cited content was actually extracted from.

Schema types present

Use Google’s Rich Results Test or view the page source to check what schema markup is implemented. The schema types that appear most consistently on cited pages are FAQPage (the strongest direct citation surface), Article or BlogPosting with named author attribution, HowTo on process content, and LocalBusiness or specific subtypes for local queries.

Record which schema types each cited page implements. After auditing ten to fifteen cited pages across your query set, you will see a clear pattern for the schema stack that characterises citation-selected content in your niche.

Content structure signals

Beyond schema, note the structural patterns visible on the page: H2 and H3 headings that correspond to specific questions rather than topic labels, numbered or bulleted lists for process content, comparison tables for evaluation queries, and FAQ sections whether schema-marked or not. These structural elements are extraction surfaces. AI systems parse structured content more reliably than flowing narrative prose.

Domain authority and topical relevance

Record the domain of each cited page and assess its authority and topical relevance. If low-authority sites with excellent content structure are being cited alongside high-authority generalist sites, content structure is the differentiating factor in that query’s citation pool. If only high-authority sites appear, domain authority is the limiting constraint regardless of content structure.

Author attribution signals

Check whether cited pages carry a named author with visible credentials. On YMYL queries, specifically health, finance, and legal, named author attribution with professional credentials is a near-consistent property of cited pages. On general informational queries it is less universal but still present on a significant share of cited sources.

Step 3: Comparing cited vs non-cited pages

A direct comparison between pages that are cited and pages that rank for the same query but are not cited produces the most actionable findings in any source analysis.

Identify two or three queries where your page or a competitor’s page ranks in the top five but is not cited in the AI Overview. Open those pages alongside the cited pages for the same query and compare them across the five dimensions above.

The typical pattern: cited pages lead with the answer while non-cited pages lead with context or framing. A page that opens with “In this article, we will explore the factors that affect how often you should see a dentist” is not extraction-ready. A page that opens with “Most adults should see a dentist every six months, though your dentist may recommend a different frequency based on your oral health history” is.

Cited pages have FAQPage schema; non-cited pages of similar quality often do not. This is the most consistent structural difference across citation pools in most niches. It is also one of the fastest gaps to close.

Cited pages use headings as questions; non-cited pages use headings as topic labels. “Factors That Affect Dental Visit Frequency” is a topic label. “How often should you see a dentist?” is a query-format heading. The second signals to extraction systems that this section answers a specific question.

Record your comparison findings in your logging spreadsheet as a separate tab: cited page URL, non-cited page URL, query, structural differences, schema differences, content structure differences.

Step 4: Extracting patterns and applying them

After three to four weeks of testing and auditing across your query set, you will have enough data to identify the specific gaps in your content and schema relative to the citation pool.

Translate the findings into a prioritised action list. The order matters.

Fix schema gaps first. If cited pages have FAQPage schema and your pages do not, that is the fastest citation-probability improvement available. Implement FAQPage schema on your target pages before addressing content structure. Schema changes deploy immediately and can influence citation within days of Google’s next crawl.

Restructure content openings second. If your pages open with framing rather than answers, edit the opening of each page to lead with a direct, self-contained answer to the query question. This is a targeted edit, not a full rewrite.

Update heading formats third. Change H2 and H3 headings on your target pages from topic labels to query-format phrasing where they do not already match.

Add FAQ sections fourth. Add a FAQ section to each target page covering the questions your source analysis log shows are being asked in the query pool for that topic. Each answer should be 50 to 80 words and directly answering the question without requiring surrounding context.

Tools that assist source analysis

Google Search Console is the starting point for identifying queries where you rank but do not get clicks. Queries with high impressions and low CTR are candidates for AI Overview investigation; they may be triggering an AI Overview that is absorbing the clicks your ranking position should be generating.

Authoritas tracks AI Overview citation patterns at scale and publishes research on citation vs ranking patterns. Their data on mid-page positions outranking position-one pages for citation is the most practically useful published research on this topic for practitioners.

Profound is the most purpose-built AI citation monitoring tool currently available. It tracks brand and domain visibility across AI engines and can surface citation data in a structured dashboard. It is primarily aimed at larger brands and agencies, but it is the closest thing to automated source analysis for practitioners who cannot run manual testing at scale.

Perplexity’s source sidebar is the most transparent AI engine for manual testing because it shows cited sources explicitly in the interface. For query-by-query source analysis, Perplexity is faster to work with than Google’s AI Overviews because you do not need to hover over or expand citations.

Browser extensions that extract metadata from a page speed up the per-page audit step. SEO Minion and Detailed let you check a page’s structured data and heading hierarchy without viewing source code.

Frequently asked questions

How do I find out which pages Google AI Overviews is citing?

The most direct method is manual testing. Open Google Search and search for your target query. If an AI Overview appears at the top of the SERP, expand it and look for the source citations, which typically appear as small links alongside the AI text or in a sources panel. Note the URLs and log them. Tools like Authoritas and Profound can automate parts of this process for larger query sets, but manual testing remains the most accurate method for query-specific source analysis.

Why does Google AI Overviews cite some pages and not others?

Citation selection consistently favours pages that lead with a self-contained, directly answering paragraph; have FAQPage, Article, or HowTo schema implemented; use headings that correspond to specific questions rather than generic topic labels; and come from domains with established topical authority for the subject matter. A page can rank at position one and still not be cited if its content is structured for narrative reading rather than AI extraction. Citation is a separate optimisation problem from ranking, and the two sets of signals only partially overlap.

Do AI Overview citations always match the top organic results?

No. Authoritas has published data showing that pages ranked in positions four through ten are regularly cited in AI Overviews when their content is better structured for extraction than the top-ranking pages. The correlation between organic ranking and citation exists but is not one-to-one. A page’s citation probability is determined by content structure, schema presence, and topical authority signals, not just by the link authority and keyword relevance signals that drive organic ranking position.

What schema types appear most often in AI Overview cited sources?

FAQPage schema is the most consistent schema type across cited pages in most niches. It is the most direct citation surface available because it makes individual question-and-answer pairs machine-readable and extractable. Article and BlogPosting schema with named author attribution are the second most common. HowTo schema appears consistently on process-type content. LocalBusiness schema and its subtypes appear on locally cited pages. A page implementing FAQPage alongside Article schema with author attribution covers the schema fundamentals for the majority of informational citation scenarios.

How often do AI Overview citation sources change?

Citation sources for a given query can change within days if a competitor publishes a better-structured page or if Google’s retrieval systems update how they weight citation signals. In practice, the citation pool for most informational queries is relatively stable month to month but shifts meaningfully quarter to quarter as competitive content landscapes evolve. Running source analysis monthly captures meaningful changes without the overhead of weekly testing that most teams cannot act on fast enough to justify.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top