ACT ONE · MEASURE · AI VISIBILITY

Five questions stand between your site and an answer that names you.

AI visibility is whether machines can fetch you, read you, cite you, and name you. Each question has a different answer, measured a different way, and the fifth — did anything move — can be observed, but never attributed. Here is what each one looks like when someone actually checks.

HOW TO READ /sites/:slug/visibility · 5 stages · formulas and measurement policy

Stage 1 runs now — robots.txt plus a live fetch for 10 AI crawlers. No signup.

PLATE 0 · THE FUNNEL STRIP · Each stage names its kind of claim beneath its value. The stages are a sequence, not a pipeline — nothing converts between them.

CHAPTER 01 · REACH

Can a machine reach you at all?

Before an answer engine can be wrong about you, it has to be able to reach you at all — and the two things that decide it disagree more often than you would think. TrustGrowth runs two checks per crawler: what your robots.txt permits, and what a live fetch under that crawler's own user agent actually returns. The second column is the one a robots.txt linter cannot give you: a CDN or WAF can block a bot that robots.txt welcomes, and it will do it silently.

Ten crawlers, twenty checks, one number. In this capture, nine of ten come back clean. The tenth is the whole finding.

SCROLL →

AI crawler access matrix: ten crawlers, each checked twice
PLATE 1 · STAGE 1 REACH · Ten crawlers, each checked twice, in this capture.
  1. 1.1 The verdict column. Two independent checks folded to one word per crawler. It does not mean nine crawlers visited; it means nine may.
  2. 1.2 The blocked row. A disallow the site's operator chose, recorded as blocked. Not an error: blocking a crawler is a decision this product reports, not reverses.

Reach: a count of ten crawlers

One observation is one named crawler: GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, PerplexityBot, Google-Extended, CCBot, Bytespider, Amazonbot, and Meta-ExternalAgent. Each gets a robots.txt permission check and a live homepage fetch under its user agent.

reachable_count = number of present crawlers, if none is unknown; otherwise null.

The API marks a crawler unknown when the site or robots.txt is unreachable, or the live fetch returns 403, 429, 500 or higher, or no status. Otherwise a robots disallow is absent; allowed and partial verdicts are present. Any unknown withholds the entire count. This is an unweighted count with ten possible crawlers, not a rate or a sample: no interval or effective sample size applies.

The latest completed readiness run must have a parseable check date and a bots list. Only robots.txt HTTP 200 and 404 are usable; other responses cannot establish permission. A capture older than 35 days is stale.

The UI currently counts verdicts other than blocked, including partial live fetches, unless the site or robots.txt is unreachable. The API withholds on any unknown crawler instead. Those counts must not be interchanged.

The separate llms.txt state never enters the count: HTTP 200 is present, 404 is absent, and an unusable response is unknown. Its body is not used in this measurement.

Run the keyless Reach check. No account is required.

Stage one costs nothing, on every plan.

CHAPTER 02 · READABLE

The stage the screen invites you to misread

The tile says 92, and under it, smaller, a label. Until this week that label said "answer-ready pages," and it invited exactly the wrong reading. It is not ninety-two pages. It is 92 out of 100 — an unweighted mean of four group scores, each a mean of its checks' pass rates over the pages the crawl could read. We changed the label. The number was never the problem; the caption was.

Readable: the formula behind the score

Readable is a 0–100 score, not a page count: an unweighted mean of four group scores, each an unweighted mean of its checks’ pass rates over eligible pages. One observation is one GEO check on one HTML page to which that check applies. Noindex pages contribute to no denominator.

check_score = 100 × (1 − min(failures, eligible) / eligible)

group_score = mean of defined check scores in the group

Readable = round(mean of defined group scores)

Structured data
missing_schema_types
Authorship / trust
missing_author_markup, missing_publish_date, missing_update_date
Answer structure
poor_heading_hierarchy, missing_faq_section, faq_not_marked_up, low_list_usage
Citation quality
low_outbound_authority_links

A check with zero eligible pages is omitted, not scored as zero. A group with no defined checks is omitted too. The incomplete_schema check is computed but belongs to no group and does not affect this score. The calculator is GeoCalculatorV3; this is not the TrustGrowth Score’s GEO pillar.

Failures are distinct pairs of check type and page URL, with trailing slashes removed, originating in the selected audit and open on the reference date. Duplicate findings count once; findings from older audits are excluded. Eligibility comes from the crawl’s geo-check-coverage-v1 manifest.

The audit must be published (complete or legacy_unpublished_contract), have a positive pages_crawled count, and a completed coverage manifest of that version. Missing coverage is not_collected, never a perfect score. Under rolling coverage, pages_crawled is the cycle batch, not the site’s full inventory.

The API selects the latest published audit with a completed run and calculates as of its audit date. The UI calculates as of today. Historical reconstruction is unavailable if a later audit has overwritten issue lineage. Freshness follows the site’s audit window.

group_coverage = number of defined groups / 4

crawl_coverage = pages_crawled / (pages_crawled + fetch_errors_count)

confidence = group_coverage × crawl_coverage

Missing fetch-error counts use crawl coverage 0.5; zero attempted pages use 0. Partial coverage lowers confidence, not the score, and yields maturing rather than measured. The calculator records partial_geo_check_or_crawl_coverage as its missing reason. The API still labels a present score measured but does not expose that reason alongside the score. Confidence is coverage, not a confidence interval or sample size. There is no n; pages_crawled is not a sample size. Readable uses the audit, not Search Console data.

92 is not ninety-two pages. It is 92 out of 100.

AI-readable pages: four group scores, sorted worst-first
PLATE 2 · STAGE 2 READABLE · Four groups, sorted worst-first, so the top bar is the thing to fix. Each is a pass rate over crawled pages, in this capture.

∎ DOES NOT MEAN

A partial crawl does not report a confident score: it downgrades to maturing, with confidence below 1.0 printed beside it. A score you cannot trust is worse than a gap you can see.

Between reading and knowing there is a screen that does something none of the others do: it prints the machine's own paraphrase of you, beside the words you chose.

How AI describes the site, next to the site's own positioning
PLATE 3 · HOW AI DESCRIBES YOU · Read by the model from the site's own pages — no search, no prior knowledge of the brand, in this capture.
  1. 3.1 The two columns. Your positioning beside the machine's reading of it. The panel renders no verdict on whether they agree; a divergence is the owner's call, not automatically a defect.
  2. 3.2 Left ambiguous. The only screen in the product that reports what a model could not work out. Not a failure state: "the model read your pages and still could not tell" is a finding about the content.

It does not tell you which column is right. That is not a gap in the feature; it is the feature declining to have an opinion about your positioning.

Curious what it reads on yours? → /tools/how-ai-sees-your-site — no account, one page, the same reader.

CHAPTER 03 · RETRIEVED

The only stage where you can read the machine's exact words

Stage three does not report that you appear in AI Overviews. It reports the sentence.

“Product observability is the practice of instrumenting a product so that feature usage, activation and drop-off are measurable in the same way infrastructure is…”

— the passage Google's AI Overview quoted from northwind.dev · rank 1 · in this capture

Among keywords your domain already ranks for, this stage is a census, not a sample. The provider returns at most 200 AI Overview reference rows per domain; a count at the cap is a floor, not a total. The count is unweighted, and it refuses to print a margin of error on a count.

Retrieved: a bounded census count

cited_keyword_count = number of unique own-domain AI Overview keywords returned.

One observation is one normalized keyword for your own domain. DataForSEO Labs ranked_keywords is queried for ai_overview_reference items with a limit of 200. Only that item type survives; duplicate keywords are collapsed case-insensitively, keeping the best rank_group. Competitor rows never enter your headline.

At 200 returned rows, the count is a floor, not a total. The run records truncated_domains; the API does not currently expose that flag. Keywords your domain does not rank for are outside this census. It is an unweighted count with no denominator, confidence interval, or n.

The latest completed retrieval run is authoritative, including a successful zero capture. Both the capture date and recorded count must exist; missing metadata is not_measured. The API reads the count recorded on that run, while the UI recounts its dated own-domain rows. After 35 days the UI withholds the headline; the API retains the count with stale: true.

Sampled AI Mode, ChatGPT and Gemini surface rates on the same screen are a separate instrument, not Retrieved. Each rate is cited answers / answered rows for that surface’s latest capture; pending and failed answers are excluded. These rates are unweighted, have a Wilson interval and a raw answered n, and do not apply Recalled’s category-intent or verified-demand filters. Surface entitlement is shown separately. Likewise, demand-weighted citation share of voice is not this census count.

SCROLL →

Stage 3 citation table: keyword, rank, and the quoted passage
PLATE 4 · STAGE 3 RETRIEVED · CENSUS · Keyword, rank, and what Google quoted — every row a real citation, in this capture.
  1. 4.1 Rank. The rank of the AI Overview reference block. Not your organic position: on the keywords screen the same keyword reads 11. Both are correct.
The gap block: keywords cited for competitors, not you
PLATE 5 · THE GAP BLOCK · Keywords where AI Overviews cite a competitor and not you, sorted by your own Search Console demand, in this capture.
  1. 5.1 The demand column. Your first-party impressions attached to each gap. Rows with no demand data sort last; they never sort as zero.

∎ DOES NOT MEAN

A capture that never ran, or whose API call failed, writes nothing at all — an empty census and a failed check are different facts, and the product keeps them different by refusing to write the second as the first.

Stage 3 needs Search Console — connect it in two clicks, read-only.

CHAPTER 04 · RECALLED

The stage where the better number loses

Stage four is the first that cannot be counted, only sampled — and it is where the product turns down the flattering figure. Sampling asks each engine, cold, with no retrieval and no system prompt: unprompted, does it name you? Branded prompts read 96% ±6. Category prompts read 53% ±11. The headline is 53%, and the reason is written into the code: branded recall is near-100% for anything that exists at all. Averaging it in would tell every site it was doing well.

SCROLL →

Recall by prompt intent and by engine, with confidence-interval bounds
PLATE 6 · STAGE 4 RECALLED · SAMPLED · Recall by prompt intent and by engine, with interval bounds drawn at reading size, in this capture.
  1. 6.1 The thin vertical rules. The confidence-interval bounds; a sampled rate and a counted one look identical without them. At card size they read as gridlines, which is why nothing here renders small.
  2. 6.2 The branded row. Near-ceiling recall, measured and shown. Not the headline, on purpose; it is true of any site that exists.
  3. 6.3 The per-engine row. Three engines, n=30 each. This capture does not distinguish the engines, and we are not going to pretend it does.
FIGURE 4c · REBUILT · The app draws these on separate rows, where the overlap is invisible. One shared axis is the honest rendering — which is why this page will not tell you that GPT knows you and Gemini doesn't. REBUILT

∎ DOES NOT MEAN

When the interval is wider than the rate, the number is withheld entirely — “sample too small to call.” Wilson, not the normal approximation, because sampling runs small and routinely produces 0-of-n, where the normal interval collapses to zero width and claims certainty.

The flattering number is right there, measured, and it is not the headline.

Four stages in, nothing has been claimed that was not counted or bounded.

The fifth stage is where every tool in this category starts lying.

CHAPTER 05 · IMPACT

Did anything move, and can we say why?

Twenty-eight days against the prior twenty-eight, branded impressions only, both windows ending three days back — because Search Console lands late, and a partial final day reads as a decline that never happened.

Stage 5 impact: branded-search window against the prior period
PLATE 7 · STAGE 5 IMPACT · CORRELATION · The caveat in the grey box is not our marketing copy. It is the product's own, on the screen, always.

FOUR WAYS THIS STAGE PRINTS NOTHING

insufficient_history
fewer than 56 days of history
no_branded_queries
no query contains your name
no_prior_baseline
a zero denominator; “+∞” is not a figure
insufficient_volume
under 30 impressions, where relative standard error exceeds ~18%

“+100%” off a base of nine impressions would be exactly the false precision this product exists to argue against. Below thirty, it prints nothing at all.

People who meet a brand in an AI answer often search for it by name afterwards. That is a reason to watch this number. It is not a reason to claim it.

WHAT THIS DOES NOT SAY

The five stages share no unit and do not compose. There is no “9 in 10 crawlers reach you and 5 of those become citations”; the chevrons between tiles are sequence, not arithmetic. And stage 5 is co-movement, never attribution: nothing on this page claims AI caused a search.

WHERE THE LINE FALLS

FREE
Stages 1 and 2 — crawler reach and answer-readiness, one site.
STARTER · $9
Adds stage 3: which of your ranking keywords Google's AI Overview cites, and the sentence it quoted.
PRO · $49
Adds stages 4 and 5: unprompted recall sampled across three engines, and the branded-search window beside it.

Frequently asked questions

What does "AI visibility" mean here?

Five separate questions, each measured a different way: can machines fetch you, can they use what they fetched, are you cited when they search, are you named unprompted, and did anything move. They are a sequence, not a pipeline. Nothing converts between them.

Can I get one AI visibility score?

No. The five stages share no unit, so there is no honest way to add them. There is no "nine of ten crawlers reach you, so five became citations" — the chevrons between the stages are sequence, not arithmetic.

How do you check whether AI crawlers can reach my site?

Two checks per crawler: what your robots.txt permits, and what a live fetch under that crawler's own user agent actually returns. The second is the one a robots.txt linter cannot give you — a CDN or WAF can block a bot that robots.txt welcomes, silently.

If robots.txt returns a 403, am I blocking AI?

No. A robots.txt 403, 429 or 5xx makes robots access unknown and withholds the Reach count. Only a readable robots.txt that disallows a crawler establishes a block. A live homepage fetch is a separate check: the UI can label it partial, while the API withholds the count on an unknown response.

Is "recalled" measured or estimated?

Sampled, and it is the only stage that carries a confidence interval for exactly that reason. When the interval comes out wider than the rate itself, the product withholds the number rather than printing one it cannot stand behind. Branded prompts are measured and shown, but kept out of the headline — near-total branded recall is true of any site that exists.

Which AI engines do you ask?

Recall is sampled across Claude, GPT and Gemini; how many of them your site samples is bounded by your plan. Citations are read from Google's AI Overviews, and that stage is a census rather than a sample — unique returned ranking keywords whose AI Overview cites your domain, with the sentence it quoted. The provider caps each domain at 200 rows, so a count at the cap is a floor, not a total.

Can you show that AI drove traffic to my site?

No, and no honest tool can. Stage five puts a branded-search window beside the prior period and calls it what it is: co-movement, never attribution. People who meet a brand in an AI answer often search for it by name afterwards — that is a reason to watch the number, not a reason to claim it.

What do I need to connect before any of this runs?

Nothing for stage one — crawler reach runs on a public URL with no account. Connecting Search Console, read-only, is what opens the later stages.

Stage one is free and needs no account.

What happens: we fetch your robots.txt and run a live per-UA fetch for ten AI crawlers — about fifteen seconds, no account. Readable comes from an audit, not Search Console. Connecting Search Console is needed for stage 3.