AI Overviews Tracker: How to Check Whether AI Can Fetch Your Site (The Visibility Chain, Part 1 — Reach)

Before an AI overviews tracker can show presence, Google must fetch your page. Learn the GSC method to verify reach, not rank or citation.

Article highlights

  • Estimated reading time: 10 minutes
  • Published on: October 4, 2026
  • Last updated: October 4, 2026
RA
Published · Updated · 10 min read
Branded cover for the article 'AI Overviews Tracker: How to Check Whether AI Can Fetch Your Site (The Visibility Chain, Part 1 — Reach)': TrustGrowth wordmark and title with callouts: Reach means confirmed crawl access, not AI Overview presence, Three states, not two: reachable, blocked, unknown, Google-Extended and Googlebot are separate directives.

Article

What this piece covers

This piece sets out how to verify whether "reachable by Google's crawlers" — the prerequisite most AI Overviews tracker dashboards silently assume — holds for a given domain, using Google Search Console's Crawl Stats report, robots.txt/user-agent checks, and the URL Inspection tool. This is a methodology piece, not a site audit: no fixed sample of URLs is scored here. We document the three data sources, their stated limits, and a five-step method any practitioner can reproduce. No ranking, citation, or AI Overview presence claim appears anywhere in this piece.

What "AI Overviews tracker" usually means (and the assumption built into it)

Most tools sold under this name — SERanking's AI Overviews Tracker, AI Clicks' tracker, Rankscale, Seobility, and Keyword.com — document a tracked-query check for whether Google's AI Overview box appears and whether a domain is cited inside it. AIOverviewTracker.com documents keyword-based appearance monitoring. Advanced Web Ranking's free Google AI Overview Tool documents how those overviews change over time and which domains and URLs they link. Those published checks are a real and useful measurement of presence, citation, or linked sources. A presence or citation result still assumes the tracked page was reachable by Google's crawlers in the first place.

If a page returns a 403, sits behind a robots.txt disallow, or has never been crawled, its absence from an AI Overview tells you nothing about content quality or relevance. It tells you the page was never in the running. Before you can debate whether a tool measured your site or just failed to reach it, you need a way to check reachability directly, independent of any tracker's output.

This piece is Part 1 of a five-part series describing what we call the Visibility Chain: Reach (Part 1) → Readable (Part 2) → Retrieved (Part 3) → Recalled (Part 4) → Impact (Part 5). Each stage is a separate, measurable question. This piece covers Reach only. TrustGrowth's broader GSC-verified scoring approach follows the same principle: verify before you score.

Reach is the first link in the chain, not a stand-in for the whole thing

Reach is a prerequisite for AI Overview inclusion, not a proxy for it. A page Google can fetch and return a 200 response for is eligible to be processed further; a page it cannot fetch is not. That is the entire claim this stage supports.

The correlation/causation distinction matters here and we state it plainly: a reachable page is not necessarily a page an AI system retrieves or cites in an AI Overview. An unreachable page is very likely an absent one, because it never entered the pipeline. These are two different statements and we keep them separate throughout this piece. Everything reported below is a fetchability signal sourced from Google Search Console and live crawler-directive checks — not AI Overview presence data, which is a Retrieved-stage or Recalled-stage question addressed in later parts of this series.

Method: three data sources, each with a named limit

GSC Crawl Stats report

The Crawl Stats report in Google Search Console breaks down requests by response code, by file type, and by Googlebot type (smartphone, desktop, image, video, and others). Google's own documentation states this report reflects crawl activity from the last 90 days and can lag behind live crawler behavior (Google Search Console Help, "Crawl Stats report," https://support.google.com/webmasters/answer/9679690). That lag sits next to any finding drawn from this report, not after it — a spike in 5xx responses shown today may reflect a server issue that has since been resolved.

robots.txt and crawler user-agent directives

Check robots.txt for directives naming Googlebot and Google-Extended specifically, since they govern different things. Google's crawler documentation (Google Search Central, "Overview of Google crawlers and fetchers," https://developers.google.com/search/docs/crawling-indexing/overview-google-crawlers) and the Robots Exclusion Protocol, RFC 9309 (IETF, published 2022-09), define what a Disallow directive signals: a request that a compliant crawler not fetch a given path. It does not signal that a non-compliant crawler will honor it, and it does not by itself confirm or deny indexing status.

State this plainly: Google-Extended governs use of content for Gemini's training and grounding, not classic Googlebot crawling for Search. Conflating the two produces a false reach verdict — a site can block Google-Extended while remaining fully crawlable by Googlebot, and vice versa is architecturally uncommon but not addressed by the same directive.

Live URL inspection

GSC's URL Inspection tool or its API endpoint returns a real-time crawled/indexed/blocked verdict for a specific URL, timestamped at the moment of the check (Google Search Console Help, "Inspect a URL," https://support.google.com/webmasters/answer/9012289). The Search Console API enforces a daily quota of 2,000 URL inspections per property (per Google's published API documentation) — state that quota in the same breath as recommending the tool. A single-URL check cannot generalize to a site's other 40,000 URLs; it verifies one path at one moment.

TrustGrowth's own GSC-verified audit runs this same Crawl Stats and URL Inspection check as part of a broader audit; the same checks are reproducible manually through the GSC UI or API without any third-party tool. For the complete procedure beyond reach alone, see the full Search Console audit playbook.

Three states, not two: reachable, blocked, unknown

A binary pass/fail model overstates what the data supports. Use three states instead:

State Definition Evidence required Reachable Confirmed crawl plus 200 response Crawl Stats entry + URL Inspection "Crawled — currently indexed" or equivalent Blocked robots.txt disallow, 4xx/5xx response, or noindex directive robots.txt rule match, or Crawl Stats/Inspection error code Unknown No crawl data yet, low-traffic page, or JS-rendering ambiguity Absence from Crawl Stats sample, or Inspection shows "URL is unknown to Google"

Units: categorical state per URL. Source: Google Search Console, Crawl Stats and URL Inspection reports.

Collapsing Unknown into Blocked or Reachable produces a confident number the underlying data does not support. A page with zero recorded crawl events in the last 90 days is not proven blocked — it may simply be low-priority in Google's crawl queue, or it may be rendered client-side in a way Crawl Stats alone cannot resolve.

Limitations: what a reach check verifies, and what it does not

A reach check verifies crawl access, HTTP response codes, robots directives, and render-blocking issues visible in Crawl Stats and URL Inspection, as of the date the check ran. That is the full scope.

It does not verify inclusion in an AI Overview, citation by an AI system, click-through rate, or ranking position. Those are Retrieved, Recalled, and Impact questions, covered in later parts of this series. We state this here, in the findings section, rather than in a closing disclaimer, because the limitation belongs beside the claim it limits.

How named AI Overview trackers document reach, as of 2026-09-24

This is a dated field note describing what each tool's own public documentation states, not a verdict on tool quality. Documentation silence on a topic is not proof the tool ignores it — we report only what is published.

  • SERanking's AI Overviews Tracker (seranking.com/ai-overviews-tracker.html) documents keyword-level presence tracking: whether a domain appears in the AI Overview box for a tracked query set.
  • AI Clicks (aiclicks.io/trackers/ai-overviews-tracker) documents a similar presence-tracking model with a stated 3-day trial window.
  • AIOverviewTracker.com documents keyword-based AI Overview appearance monitoring.
  • Rankscale (rankscale.ai) documents visibility tracking across AI Overviews and conversational AI systems including ChatGPT.
  • Advanced Web Ranking's Google AI Overview Tool (advancedwebranking.com/free-seo-tools/google-ai-overview) documents a free view of how AI Overviews change over time, with domains and URLs linked from AI Overviews, grouped by industry and intent.
  • Seobility (seobility.net/en/ai-overview-tracker/) documents keyword-level checks for whether an AI Overview appears and whether a site is cited, including the cited URLs and their order.
  • Keyword.com publishes a tracking-methods page (keyword.com/blog/how-to-track-ai-overviews/, updated 2026-07-27) describing keyword-set checks for AI Overview appearance and citation.
  • Ahrefs publishes a methodology for tracking AI Overview mentions and citations at scale (Ahrefs, "How to Track AI Overviews," https://ahrefs.com/blog/how-to-track-ai-overviews), documenting a citation-tracking approach as one example among several.

Our position is a choice, not a claim about any competitor's rigor: we verify GSC-confirmed reach before reporting anything about AI Overview presence, because a domain absent from a tracker's results could be genuinely unselected by Google's ranking systems, or it could have never been successfully fetched. A crawl-status check is what distinguishes those two cases for a given site — the tracker output alone cannot.

Run your own reach check: a five-step method

  1. Pull GSC Crawl Stats filtered by Googlebot type (smartphone, desktop, image) for the last 90 days.
  2. Check robots.txt for Googlebot and Google-Extended directives, noting any path-level Disallow rules.
  3. Run URL Inspection on a defined sample — state the sample size and date, for example "40 URLs, inspected 2026-08-18."
  4. Log response codes and inspection timestamps for each URL in a spreadsheet or database.
  5. Classify each URL as Reachable, Blocked, or Unknown using the three-state model above.
# Example: fetch robots.txt and grep for relevant user-agents
curl -s https://example.com/robots.txt | grep -A3 -i "user-agent: googlebot\|user-agent: google-extended"

A Blocked or Unknown result does not mean "never cited." Citation is a Retrieved-stage question this method cannot answer — it only tells you whether the page was fetchable, not what happened to it after fetch. For the full crawler-access procedure, see check whether AI crawlers can access your site.

Reading your own results: sample size and data window

A spot check of 20 or 40 URLs is not a site-wide census for a domain with 10,000+ indexed pages. State the sample size (n) next to any classification summary — "32 of 40 sampled URLs classified Reachable, sampled 2026-08-18" is a defensible statement; "our site is reachable" is not.

Crawl Stats and URL Inspection each report against Google's own stated data windows (90 days for Crawl Stats, real-time for URL Inspection). Link to the source documentation rather than restating the window from memory, since Google can change it. Keep "reach was verified on N URLs as of [date]" strictly separate from any statement about whether a site is "AI-visible," a claim this method does not support.

FAQ

How to track AI Overviews?
Use a presence-tracking tool such as SERanking's AI Overviews Tracker, Rankscale, or Ahrefs' documented methodology (ahrefs.com/blog/how-to-track-ai-overviews) against a fixed keyword set, and re-run it on a consistent schedule since AI Overview presence changes per query and per session.

Can I remove AI Overviews?
Google does not offer a universal opt-out for AI Overviews as a search feature; content owners can influence inclusion of their own pages through crawlability, structured data, and content signals, but cannot remove the AI Overview box itself for a query.

How do I get AI Overviews?
A page must first be reachable (this piece), then readable and retrievable by Google's systems (Parts 2 and 3 of this series) before it can be recalled and cited. Reach alone does not cause inclusion.

Should I trust AI Overviews?
AI Overviews should be evaluated the same way any generated summary is: check the cited sources directly, since generated text can misstate or oversimplify source material.

What is the 30% rule in AI?
There is no single documented standard called the "30% rule" that governs AI Overview inclusion or AI system behavior. Google's own documentation on AI features in Search names no percentage threshold of any kind. Treat any claim of a fixed percentage threshold as unverified unless it cites a named, dated source.

How often are AI Overviews wrong?
We have not run an accuracy audit of AI Overview content and make no claim about error rate here; that is outside the scope of a reach measurement.

Has AI ever made a mistake?
Yes, generated systems including AI Overviews have produced documented factual errors reported in news coverage, but a general error-rate figure is outside what this reach-focused method measures.

Why is AI getting worse?
We have not measured a directional trend in AI system quality and do not make a claim on this question; it is not something a Reach check can answer.

What is more accurate, Google or AI?
We have not run a comparative accuracy study between traditional Google Search results and AI Overview output and make no claim here.

Key takeaways

  • Reach means confirmed crawl access (a 200 response plus a Crawl Stats or URL Inspection record), not AI Overview presence.
  • Three states, not two: Reachable, Blocked, and Unknown, sourced from GSC Crawl Stats (90-day window), robots.txt/user-agent checks, and URL Inspection (real-time, 2,000/day API quota).
  • Google-Extended and Googlebot are governed by separate directives; conflating them produces a false reach verdict.
  • Named AI Overview trackers — SERanking, AI Clicks, AIOverviewTracker.com, Rankscale, Advanced Web Ranking, Seobility, Keyword.com, and Ahrefs' documented method — document presence, citation, or linked-source tracking on their published pages.
  • This piece documents a reach-checking method only, using GSC Crawl Stats, robots.txt/user-agent checks, and URL Inspection. It reports no measurement of its own, and makes no ranking, citation, or traffic claim.

What comes next in the Visibility Chain

Reach is the prerequisite, not the finding most readers want. The next stage, Readable, asks whether a reachable page is structured so an AI system can parse it once fetched — covered in Part 2, and related to the technical SEO archive on this site.

Later stages, Retrieved and Recalled, ask whether AI systems actually pull from and remember a page; Impact asks whether that translates into measurable demand. Each of those stages requires its own named, dated method, published separately as part of TrustGrowth's AI visibility diagnostics. None of those claims belong in a Reach measurement.

The TrustGrowth proof page separately reports a published score for trustgrowth.ai, dated on that page. That score is an unrelated, separately measured artefact. It is not evidence about reach, it says nothing about crawl access, and we do not use it here.

This piece documents a reach-checking method only, using GSC Crawl Stats, robots.txt/user-agent checks, and URL Inspection. No run of it is reported here, and no ranking, citation, or traffic claim is made.

technical SEO Google Search Console E-E-A-T gsc audit AI visibility AI Overviews crawlability
Share:

Know your site's real SEO score

Free GSC-verified audit, E-E-A-T scoring, and AI-powered content strategy.

Get Started Free

Related Articles