Why AI Visibility Scores Differ: Two Tools, One Site, 0% and 33% — and Both Were Right

Why do two AI visibility tools score the same site 0% and 33%? Prompt taxonomy, sample size, model scope, run date, and result-state handling explain the difference.

Article highlights

  • Estimated reading time: 10 minutes
  • Published on: September 21, 2026
  • Last updated: September 21, 2026
RA
Published · Updated · 10 min read
Branded cover for the article 'Why AI Visibility Scores Differ: Two Tools, One Site, 0% and 33% — and Both Were Right': TrustGrowth wordmark and title with callouts: 0% and 33% can both be right on the same site, Six dimensions: taxonomy, sample, models, date, non-answers, aggregation, Report recall and discovery separately, per model.

Article

Abstract

What this piece measures: how two AI-visibility instruments can score the same site 0% and 33% recall while both remain methodologically defensible, once you account for prompt taxonomy, sample size, model scope, and result-state handling. The worked example uses hypothetical, explicitly labeled figures to demonstrate the arithmetic; it is not a live audit of a named site. Every real figure carries a named source; every constructed figure is marked hypothetical. The goal is to explain measurement divergence, not to rank instruments or sites.

If you audit sites for a living, you have seen this pattern: two vendors run an "AI visibility" check on the same domain and return numbers that do not just differ, they contradict. One says the site is invisible to generative answer engines; another says it shows up a third of the time. Neither number is fabricated. They are measuring different constructs and calling them the same thing.

Method

Before any figure means anything, the measurement design has to be stated. The comparison below is organized around six dimensions that determine what an AI-visibility score actually represents:

  • Prompt taxonomy: whether prompts are branded (naming the company or product), unbranded (describing the need or category without a brand name), or a mix, and how that classification is defined.
  • Sample composition: the ratio of branded to unbranded prompts, and whether it matches how real users query assistants like ChatGPT, Claude, or Gemini for the topic.
  • Sample size: the raw count of prompts run, since a 12-prompt and a 400-prompt sample carry very different confidence intervals even at the same percentage.
  • Model coverage: which models were queried (GPT-4o from OpenAI, Claude 3.5 Sonnet from Anthropic, Gemini 1.5 Pro from Google), and whether results are reported per model or blended.
  • Run date: the calendar date(s) prompts were executed, because model weights, retrieval layers, and grounding sources change over time and are not versioned publicly by most vendors.
  • Result-state handling: how a tool codes an answer that does not mention the site — as absent, blocked, no-data, or excluded from the denominator.

Any fact about a specific product, model, or methodology below carries a named source. Where no dated primary source exists for a figure, it is labeled hypothetical and used only to illustrate arithmetic, consistent with how we document our own AI visibility measurement: we report recall per model with a confidence interval and record a non-mention as its own result-state rather than forcing it to a zero.

Why AI visibility scores differ

Changing one measurement dimension changes the construct being measured, even when the output is still labeled "AI visibility score." A 12-prompt, unbranded-only, single-model sample answers a narrower question ("does GPT-4o surface this site for category questions") than a 40-prompt, blended, three-model sample answering "across the three most-benchmarked assistants, how often does this site appear for any relevant query." Both are legitimate. They are not the same measurement.

Recall and discovery are also frequently conflated. Recall asks whether a model mentions or cites a specific site when prompted on a relevant topic. Discovery asks whether a model can find and correctly describe the site at all, independent of ranking against competitors. A tool reporting "33% AI visibility" without specifying which it measured has given you a label, not a comparable number.

Per-model results and blended summaries diverge for a related reason: models retrieve and generate differently. Claude, GPT-4o, and Gemini do not share a retrieval index, and public documentation on each model's grounding sources is limited, so a site cited by one model and absent from another is expected behavior of three separate systems, not an anomaly. Blending three models into one percentage hides which model is driving the number.

Applied case study: one site, different measurement designs

The table shows how two differently designed instruments (Instrument A and Instrument B, to avoid naming a competitor) could report 0% and 33% for the same site. All figures are hypothetical and constructed for this article; they are not a measured audit of any named site.

Dimension Instrument A Instrument B Prompt taxonomy Unbranded-only (category questions) Blended: 70% unbranded, 30% branded Sample size 12 prompts 40 prompts Model coverage 1 model (GPT-4o) 3 models (GPT-4o, Claude 3.5 Sonnet, Gemini 1.5 Pro) Result-state handling "No mention" counted as absent, in denominator "No mention" excluded from denominator Run date 2025-01-14 2025-02-03 Reported score 0% recall 33% recall

Units: percent recall (0–100%); prompt count (n); model count; date (ISO 8601). Source: hypothetical figures constructed to demonstrate how measurement design changes the reported score; not a measured run on a named site.

Instrument A's 0% is true about a narrow construct: zero mentions across 12 unbranded prompts on one model, on one date. Instrument B's 33% is also true about a broader construct: one mention in three across 40 blended prompts on three models, with absent results excluded from the denominator. Blended score B is a simple mention rate over its included denominator, not a weighted composite. Neither vendor lied; they answered different questions and printed the same label on both answers.

To confirm the table is internally consistent, we recomputed it by hand: Instrument A's denominator is all 12 prompts with zero mentions, so 0 ÷ 12 = 0%. Instrument B excludes non-answers from its denominator and records roughly one mention for every three counted responses, so its rate rounds to 33% (1 ÷ 3). The two figures are arithmetically sound for the two designs described — they remain constructed illustrations, not a measured audit of any site.

Authority signals are a separate axis

Treating backlink or authority data as part of the AI-visibility number is a related confusion. In a hypothetical example, a site with 39 referring domains (illustrative, not measured) tells you nothing directly about how often that site is cited inside an AI-generated answer — authority and AI recall are measured on separate axes and should be reported separately, which is why authority and backlink analysis is scored independently of AI visibility.

We apply the same separation-of-constructs discipline to our own scoring, and we publish the results even when a sub-score comes in low: when our own authoritativeness scored 33 (published August 6, 2026) we reported it on a documented rubric rather than an unexplained figure. We publish public proof pages alongside a score so a reader can check which prompts, models, and dates produced the number.

Prompt taxonomy and sample composition

Prompt taxonomy — the branded/unbranded split — is one input to measurement design, not the whole method, and this piece treats it as a single dimension among the six rather than re-teaching the split itself. What matters for divergence is that the branded/unbranded ratio is a methodology choice with measurable effects, and independent GEO research treats the split as foundational to interpreting any AI-citation dataset — see Britopian's research on branded vs unbranded prompts in generative AI search (published 2025-07-24) and DeepSmith's analysis of AI citations vs mentions (published 2026-07-09).

Sample composition changes interpretation even when the taxonomy definitions match: a sample skewed heavily toward branded prompts scores higher recall for an established brand than an unbranded-heavy sample, so a tool that does not disclose its branded/unbranded ratio is not comparable to one that does. maxaeo.ai's measurement guide on branded vs non-branded prompts (published 2026-07-14) makes the same point: the ratio is a methodology choice, not a neutral default.

A minimal, reproducible taxonomy config might look like this:

{
 "prompt_set": "unbranded_category",
 "sample_size": 40,
 "branded_ratio": 0.30,
 "models": ["gpt-4o", "claude-3.5-sonnet", "gemini-1.5-pro"],
 "result_states": ["cited", "mentioned", "absent", "blocked"],
 "run_date": "2025-02-03"
}

Any AI-visibility report that omits an equivalent configuration is asking you to trust a percentage without the sample it was drawn from.

Model scope and test date

Model selection bounds what a score can claim. A single-model test (GPT-4o only) tells you about GPT-4o's behavior on that date, nothing about Claude or Gemini. A three-model test tells you about the blend of three systems, which requires its own aggregation rule stated explicitly, whether a simple average, a majority vote, or a weighted composite.

Dated runs matter because no major provider publishes a stable versioning guarantee for consumer-facing retrieval behavior. A recall percentage from a run dated 2025-01-14 describes that date's model weights and retrieval layer, not a permanent property of the site or the model. Treat every AI-visibility figure as a time-bounded sample, the same way you would treat a Google Search Console query pull for a fixed date range rather than a permanent ranking fact.

How to compare AI-visibility measurements responsibly

Before comparing two AI-visibility scores, check that each report discloses:

  1. Taxonomy: how branded and unbranded prompts are defined and classified.
  2. Sample count and composition: total prompt count (n) and the branded/unbranded ratio.
  3. Models: which named models were queried and whether results are per-model or blended.
  4. Date: the ISO 8601 date(s) the prompts were run.
  5. Unavailable/blocked handling: whether a non-answer counts as absent, is excluded, or is flagged separately.
  6. Aggregation rule: the exact formula that turns per-prompt results into one headline number.
  7. Limitations: what the sample does not cover and why.

For how one scoring system documents these choices, see how the TrustGrowth score is calculated. Any score without a visible method attached should be treated as a label, not a measurement.

Limitations

This article uses a single constructed example to illustrate arithmetic, and that example does not generalize to any real site's actual AI visibility. The sampling error at these sizes is large by definition: using the Wilson score interval for a binomial proportion from the NIST/SEMATECH e-Handbook of Statistical Methods' confidence-interval guidance for proportions (published 2003, last updated 2012), a 12- or 40-prompt sample cannot support a 95% interval much narrower than roughly ±15 to ±20 percentage points near a mid-range mention rate, so treat any small-sample AI-visibility score as directional, not precise.

Measurement differences documented here explain why two numbers diverge; they do not establish which site is higher quality, more authoritative, or more likely to convert. This piece makes no ranking, traffic, or revenue claim about any site, tool, or model. Every product, model, and methodology reference above is time-bound: vendors change taxonomies, providers update retrieval behavior, and a configuration accurate on its run date may not hold today.

FAQ

What actually determines what an AI-visibility score measures?

Prompt taxonomy, sample composition, sample size, model coverage, run date, and result-state handling each change the construct behind the number. The branded/unbranded split is only one of the six; a score reported without all six disclosed is a label, not a comparable measurement.

Why can two tools score the same site 0% and 33% differently?

Because each measures a different construct: differences in prompt taxonomy, sample size, model coverage, run date, and how non-answers are counted can move a reported score by tens of points while both tools remain internally correct. A 0% and a 33% can both be true statements about different measurement designs.

Should AI visibility be tested on more than one model?

Yes. Claude, GPT-4o, and Gemini do not share a retrieval index and can return different answers to an identical prompt on the same date, so a score built on one model cannot be generalized to "AI search" as a whole. Report per-model results before any blended figure.

How should a non-answer be counted?

Consistently, and disclosed. Whether a "no mention" is coded as absent (in the denominator), excluded from the denominator, or flagged separately as its own state materially changes the percentage, so the choice must be stated for two scores to be comparable.

What is an example of a branded prompt?

A branded prompt names the company or product directly — for example, "Is TrustGrowth good for GSC-verified audits?" The unbranded version of the same intent removes the name and describes the need: "best tool to audit E-E-A-T signals for a SaaS site." The first tests whether a model already recalls the brand; the second tests whether it surfaces the brand competitively when no name is given.

Is a branded keyword the same as a branded prompt?

They overlap but are not identical. A branded keyword is a search query typed into a search engine that contains the brand name, returning a ranked page of results. A branded prompt is a natural-language request to an assistant like ChatGPT, Claude, or Gemini that names the brand and is answered by a single synthesized response. Both include the brand token, but they are measured differently: a keyword by rank position, a prompt by whether the one answer mentions the brand at all.

Key takeaways

  • A 0% and a 33% AI-visibility score on the same site are not necessarily contradictory once you check prompt taxonomy, sample size, model coverage, run date, and result-state handling.
  • Branded vs unbranded prompt composition is one of at least six measurement dimensions, never the whole methodology by itself.
  • Recall (does the model cite this site) and discovery (can the model find it at all) are different constructs and should be reported separately, per model where possible.
  • Every AI-visibility figure is a time-bounded sample tied to a specific run date, not a permanent property of a site.
  • Before comparing two scores, require disclosure of taxonomy, sample count and composition, model coverage, run date, non-answer handling, and the aggregation formula.
E-E-A-T GEO branded vs unbranded ai prompts AI visibility measurement AI search citations prompt taxonomy measurement methodology
Share:

Know your site's real SEO score

Free GSC-verified audit, E-E-A-T scoring, and AI-powered content strategy.

Get Started Free

Related Articles