Sitemap Discovery Intelligence
Audit freshness signals, structural problems and priority-page coverage—then compare sitemap declarations with observed AI crawler activity.
Discovery readiness report
The Sitemap Freshness & AI Discovery Auditor reviews sitemap URL coverage, last-modified signals, structural problems, priority-page inclusion and optional crawler evidence so you can identify URLs that may be poorly represented in your discovery infrastructure.
An XML sitemap tells search engines about URLs and files you consider important and can provide additional information such as when a page was last significantly updated. It helps discovery and crawl planning, but inclusion in a sitemap does not guarantee that a URL will be crawled, indexed, ranked or cited by an AI system.
A sitemap is often treated as a set-and-forget WordPress file.
That misses its real diagnostic value.
A useful sitemap should contain the current preferred URLs you want search systems to discover, use accurate modification dates when available and stay aligned with the site's redirects, canonicals and publishing activity.
This auditor examines those signals rather than merely confirming that the XML parses.
A URL can appear perfectly in an XML sitemap and still remain uncrawled, excluded, canonicalized elsewhere, blocked, low value or otherwise absent from Google's index.
It tells a search system which URLs you consider important. The system still decides whether, when and how those URLs are crawled, indexed and served.
Add one or more XML files containing actual page-level URL entries.
This provides context for detecting URLs that appear outside the intended site or hostname.
Use a timeframe appropriate to the content type rather than assuming every old lastmod is a problem.
Enter commercially or editorially important pages that should receive explicit sitemap coverage.
Add crawler-log or URL-status CSV evidence when you want to compare sitemap declarations with observed requests, status codes or canonicals.
Inspect freshness, URL issues, discovery gaps and action priorities before changing the sitemap generator.
<sitemapindex>
<sitemap>
<loc>https://example.com/post-sitemap.xml</loc>
</sitemap>
<sitemap>
<loc>https://example.com/page-sitemap.xml</loc>
</sitemap>
</sitemapindex>
<urlset>
<url>
<loc>https://example.com/example-page/</loc>
<lastmod>2026-08-20</lastmod>
</url>
</urlset>
Google's current guidance says to include the URLs you want to see in Search and generally list the preferred canonical version rather than every duplicate URL that exposes the same content.
Prefer the URL version your site considers the primary representative.
Avoid listing obsolete HTTP URLs when HTTPS is the intended current destination.
Remove retired URLs when they are no longer intended as search destinations.
Yes. Google lists sitemap presence among the signals its systems can consider during canonicalization. However, canonicalization uses multiple signals, and Google may still select a different URL than the version listed in your sitemap.
Lists the preferred URL.
Expresses another explicit canonical preference.
Can communicate that an old URL has moved to another destination.
Listing URL A while URL A redirects or canonicalizes to URL B sends a less consistent set of technical signals than simply listing URL B.
Google says it uses <lastmod> when the value is consistently and verifiably accurate. It should reflect the date or time of the page's last significant update.
Modification dates change when the page itself receives a meaningful update.
The declared date is reasonably consistent with observable page changes.
The site does not constantly claim fresh modifications for unchanged URLs.
A truthful older date is more useful than an automatically updated date that does not correspond to a meaningful page change.
A mathematical definition or stable historical fact may remain accurate for years.
Software documentation, regulations, platform features and pricing can require much more frequent review.
Freshness expectations should follow the real-world topic, not one sitewide age threshold.
| XML field | Current Google behavior |
|---|---|
<loc> |
Essential URL location information. |
<lastmod> |
Can be used when consistently and verifiably accurate. |
<priority> |
Ignored by Google. |
<changefreq> |
Ignored by Google. |
If your SEO plugin still outputs these values, do not mistake them for current Google ranking or crawl-priority controls.
<loc>https://example.com/services/local-seo/</loc>
<loc>/services/local-seo/</loc>
Google's current sitemap guidance limits an individual sitemap to 50,000 URLs or 50 MB uncompressed. Larger inventories should be divided into multiple sitemap files and can be organized through a sitemap index.
Maximum 50,000 URLs per sitemap.
Maximum 50 MB when uncompressed.
Use an index to organize multiple child sitemap files.
No. Google's sitemap documentation explicitly says the order of URLs in a sitemap does not matter.
Strong internal architecture, useful content, prominent navigation, relevant links and consistent technical signals are more meaningful than manually rearranging XML entries.
Important services, categories, products and lead-generation pages.
Major guides, research pieces, tools and resources that support the site's expertise.
Newly published pages you want to monitor closely during discovery and indexing.
Review why a strategically important indexable page is absent from the sitemap generator.
Search engines may discover the URL through internal links, external links, feeds, redirects or previously known crawl paths.
But if the page is important enough to be a preferred Search destination, sitemap omission is worth understanding.
A sitemap can provide a discovery path to a URL that is otherwise poorly linked. But important pages should normally participate in sensible internal linking so users and crawlers can reach them through the site's architecture as well.
If a commercially important page is included in XML but has no contextual or navigational relationship with the rest of the site, investigate the internal-linking problem separately.
Sitemap → old URL → 301 → new URL
Sitemap → current final canonical URL
| Observed status | Potential sitemap action |
|---|---|
| 200 | Normal result for many current indexable page URLs. |
| 3xx | Review whether the sitemap should instead list the final destination. |
| 404 / 410 | Remove obsolete sitemap entries unless a temporary migration workflow explains them. |
| 5xx | Investigate server availability rather than assuming the sitemap itself is the problem. |
Sitemap URL:
https://example.com/seo-guide/
Declared canonical:
https://example.com/seo-guide/
Sitemap URL:
https://example.com/seo-guide/?source=nav
Declared canonical:
https://example.com/seo-guide/
The live auditor can compare sitemap URLs with previously observed crawler-request evidence. That can help reveal priority URLs that were declared in XML but were not observed in the supplied log window, as well as URLs crawlers reached even though they were missing from the sitemap.
| Sitemap? | Observed crawler request? | Interpretation |
|---|---|---|
| Yes | Yes | The URL was declared and also observed in the supplied request evidence. |
| Yes | No | No request was observed in the evidence window. This does not prove the crawler cannot access the URL. |
| No | Yes | The crawler discovered the URL through another route or knew it previously. |
| No | No | There is little discovery evidence in the supplied datasets; investigate further if the URL is strategically important. |
User-agent strings can be imitated. If crawler identity materially affects your analysis, use the relevant provider's official verification method where one exists rather than treating the user-agent label alone as proof.
A crawler request does not prove indexing, model ingestion, answer selection, citation, ranking or user visibility.
No. Google's current guidance says the normal way Search finds and processes pages remains the foundation for its generative AI features. A page must meet Google's normal Search technical requirements and be eligible for Search; there is no separate sitemap format required for AI Overviews or AI Mode.
Google must be able to access the content under the normal Search crawling rules.
Google says pages need normal Search eligibility to appear in its generative Search experiences.
Google explicitly says special AI markup and files such as llms.txt are not required for Google Search.
“This URL was explicitly listed in the supplied discovery file.”
“A request matching the supplied crawler evidence was observed for this URL.”
Citation and answer visibility require separate measurement. The sitemap auditor deliberately stops at the discovery and crawler-evidence layer.
Sitemap submission is only one discovery and canonicalization signal. Google still evaluates crawl access, HTTP behavior, canonicalization, duplication, page quality and other indexing conditions before deciding what to include.
The page explicitly requests exclusion from Search.
Google may treat another URL as the representative version.
The URL may belong to a duplicate cluster rather than appear separately.
Use sitemap analysis for discovery signals and a dedicated indexability/index-evidence workflow for Google's actual indexing state.
Search Console can provide information about whether Google could fetch and process a submitted sitemap.
Use Search Console reporting to investigate sitemap parsing or submission problems.
Combine sitemap reporting with page indexing and URL Inspection evidence rather than assuming submission equals indexing.
Sitemap: https://example.com/sitemap_index.xml
Identify the current replacement for every important historical address.
Route old URLs to their relevant new destinations.
List the new preferred URLs rather than perpetuating obsolete destinations.
Point the live site directly to the new URLs.
Use Search Console to observe migration processing and URL discovery.
Verify that old hosts, protocols and slugs no longer dominate current sitemap files.
Confirm that intended posts, pages, products and other public content types appear in the correct sitemap files.
Review whether categories, tags or custom taxonomies belong in Search before including them automatically.
Make sure the plugin's modification dates reflect meaningful content changes rather than unrelated template activity.
Core WordPress, an SEO plugin and a separate sitemap plugin can create overlapping discovery files. Choose one deliberate sitemap system where practical.
| Content type | Freshness concern |
|---|---|
| Platform feature guide | High when interfaces, rules or capabilities change frequently. |
| Legal / regulatory guide | High when rules or deadlines change. |
| Pricing page | High when prices, plans or availability change. |
| Evergreen definition | Lower when the underlying fact remains stable. |
| Historical article | May intentionally retain its original information and context. |
| Product specification | Review whenever the product or model changes. |
An important indexable URL is absent from the supplied sitemap set.
The sitemap still contains a historical URL that redirects elsewhere.
An obsolete URL remains in XML after the content has disappeared.
The listed URL declares a different preferred canonical destination.
Every URL receives today's lastmod despite no meaningful page update.
A frequently updated site provides no modification signal even though reliable dates are available.
Old domains, staging hosts or incorrect protocol variants remain in current sitemap files.
The auditor has sitemap-file locations but no actual page-level URL entries to analyze.
A crawler request or sitemap listing is incorrectly treated as proof of AI answer visibility.
Make sure the destination is still strategically useful and intended for Search.
Verify that the final page is accessible and does not redirect unexpectedly.
Review robots, noindex and other technical eligibility controls.
Confirm that the page is not intentionally or accidentally canonicalized elsewhere.
Make sure the page participates in the site's normal architecture rather than existing only in XML.
Use Search Console and server logs when you need stronger evidence about Google discovery, crawling or indexing.
Use canonical current destinations instead of duplicate and historical URL variants.
Include the complete protocol, hostname and path.
Update it when a meaningful page change occurs rather than whenever the sitemap is rebuilt.
Periodically confirm that high-value URLs remain represented in the current sitemap set.
Redirects, deleted URLs and deprecated variants should not silently accumulate forever.
Sitemaps, redirects, internal links and canonical annotations should generally point toward the same preferred URLs.
Use first-party reporting to monitor Google's ability to fetch and process important sitemaps.
Interpret log activity as observed requests rather than proof of indexing, citation or ranking.
Everything after discovery—crawl timing, processing, canonicalization, indexing, ranking and citation—requires additional evidence.
Direct answers to common questions about sitemap URLs, lastmod, indexing, canonicalization, Search Console and AI crawler evidence.
It is a machine-readable file that provides search engines with information about pages and other files you consider important on your site.
Not necessarily. Google says it can usually discover a small, comprehensively linked site without one, although most sites can still benefit from a sitemap.
Yes. Google says sitemaps can improve URL discovery and crawling, especially for large, new or complex websites.
No. Google explicitly describes sitemap submission as a hint and does not guarantee that every listed URL will be crawled.
No. Discovery, crawling and indexing are separate stages.
No. A sitemap is primarily a discovery and crawl-support mechanism, not a ranking guarantee.
Google recommends including URLs you want to see in Search and generally using the preferred canonical versions.
Current ongoing sitemaps should generally list the final preferred destinations rather than unnecessary redirect sources.
Normally no. Remove obsolete URLs that are no longer intended as current Search destinations.
Usually not in a sitemap intended to represent the URLs you want shown in Google Search, because the noindex instruction expresses the opposite indexing intent.
It is an XML sitemap field that communicates when the page was last significantly modified.
Yes, when Google finds that the values are consistently and verifiably accurate.
No. It should change when the page receives a significant modification, not merely because a new day began or the sitemap was regenerated.
Google specifically gives copyright-date changes as an example of something that is not considered a significant content update.
Yes. Google's sitemap guidance says meaningful structured-data changes can be considered significant modifications.
Yes. Google also lists link changes among the types of updates that may be significant.
No. Evergreen content can remain accurate for years. Freshness should be judged in the context of the subject matter.
Choose one based on the rate at which the subject realistically changes. The live tool's threshold is an editorial review setting, not a Google rule.
No. Google's current documentation says it ignores the sitemap changefreq value.
No. Google currently ignores the XML sitemap priority field.
No. Google explicitly says the ordering of URLs within the sitemap does not matter.
Google currently limits an individual sitemap to 50,000 URLs.
Google currently limits each sitemap to 50 MB uncompressed.
Split the inventory into multiple child sitemaps and organize them through a sitemap index if convenient.
It is an XML file listing multiple sitemap files so search engines can discover and process a larger sitemap collection.
The sitemap index contains sitemap locations rather than the individual page URLs needed for page-level freshness and discovery analysis.
Google recommends fully qualified absolute URLs, including protocol and hostname.
Yes. Google lists sitemap inclusion among the signals it can consider when choosing a canonical URL.
No. Google can choose a different canonical after evaluating its full set of canonicalization signals.
They generally should support the same preferred-URL strategy unless you have a deliberate technical reason for something different.
Yes. Google can discover URLs from internal links, external links and other crawl paths. Sitemap omission does not automatically make a page invisible.
It can provide a discovery path, but important pages should normally also be integrated into sensible internal linking.
Yes. Google supports Sitemap declarations in robots.txt and allows multiple sitemap lines.
Search Console is still useful because it provides first-party sitemap processing and Search diagnostics that robots.txt does not provide.
It is a user-defined list of strategically important pages whose sitemap and crawler coverage you want to check. It is unrelated to the XML priority field Google ignores.
It is a discrepancy in the supplied evidence, such as an important URL missing from the sitemap or a sitemap URL with no observed crawler activity in the imported evidence window.
No. It only means the request was not present in the supplied evidence set or time period.
No. A server request proves only that the request occurred in the supplied evidence. Citation and answer visibility are separate events.
Yes. User-agent text alone does not provide cryptographic proof of crawler identity.
No. Google's current guidance says normal Search crawling, indexing and technical eligibility remain the foundation for its generative AI features.
No. Google does not document a separate sitemap format for AI Overviews.
No. Google's current generative AI Search guidance says Google Search does not use llms.txt as a special visibility mechanism.
No. Sitemap inclusion does not guarantee crawling, indexing, Search visibility or generative AI inclusion.
No. Use Search Console or a dedicated Google Indexability & Index Evidence workflow for that question.
No. The live auditor evaluates supplied sitemap and optional CSV evidence rather than fetching every live robots.txt file or page.
Not directly. Server and firewall behavior requires live request or log evidence.
No. The live tool states that sitemap files and optional CSV evidence are processed locally in the browser.
The live tool intentionally avoids remote sitemap fetching because browser CORS restrictions and privacy constraints make direct cross-site fetching unreliable.
No. A cleaner sitemap can improve discovery communication and technical consistency, but rankings depend on many additional content, quality, relevance, authority and competitive factors.
Use the Sitemap Freshness & AI Discovery Auditor to find missing priority URLs, stale or unreliable lastmod signals, canonical conflicts and discovery gaps—then validate crawler, indexing and AI visibility with the appropriate evidence instead of treating XML inclusion as proof.