Branded vs Unbranded AI Prompts: The Split That Decides Your Score
Branded vs unbranded AI prompts measure recall vs discovery. See why the ratio in your prompt set decides your AI visibility score before a single prompt runs.
Article highlights
- Estimated reading time: 9 minutes
- Published on: September 7, 2026
- Last updated: September 7, 2026
Article
Abstract
What we measured: how the ratio of branded to unbranded prompts in an AI visibility test changes the resulting score. Method: a documented prompt-design framework, run per model (Claude, GPT, Gemini) and timestamped per run; the percentages below are labeled hypothetical illustrations, not a dated measurement of any named site. Takeaway: branded prompts test whether a model already recalls you; unbranded prompts test whether a stranger would discover you — and averaging them into one unlabeled number hides which one you measured. No outcome, ranking, or revenue claim is made.
Method
This piece describes a reproducible prompt-design method, not a new dataset. The parameters are:
- Data source: branded prompts are pulled from a site's real Google Search Console branded queries; unbranded prompts are hand-written as category-language questions in the buyer's own vocabulary.
- Buckets: branded and unbranded are kept as two separately sourced lists, never one mixed set.
- Models: Claude, GPT, and Gemini, each queried independently on a stated date.
- Reporting: results are timestamped per run and reported per model, never averaged into one hidden cross-model blend.
- Sample disclosure: each bucket's size is disclosed alongside its score. Where a percentage appears below (for example, a "70–80% branded" mix), it is a hypothetical illustration of how ratio affects a score, not a measured result for any site.
A worked, dated application to a real named site is out of scope here and handled separately.
What branded and unbranded actually mean in an AI visibility test
A branded prompt names the brand or product directly: "Is TrustGrowth good for GSC audits?" or "What does Ahrefs' Site Audit check?" An unbranded prompt describes a problem or category with no brand name attached: "best tool to audit E-E-A-T for a SaaS site" or "how do I check if my site is indexable by AI crawlers."
This mirrors a split used in keyword research for years, where branded search volume is separated from non-branded demand to avoid inflating a site's apparent organic pull. AI chat answers behave differently from a search results page, though. A Google search results page returns ten or more ranked URLs a user can scan and compare. Claude, GPT, and Gemini return one synthesized answer, so there's no equivalent of "ranked eighth": a brand is either present in that single answer, absent from it, or the model can't reach the site at all.
Branded prompt example Unbranded prompt example What it tests "Is TrustGrowth good for GSC-verified audits?" "Best tool to audit E-E-A-T signals for a SaaS site" Whether the model recalls a known name vs. surfaces the site competitively with no name given "What does the TrustGrowth score measure?" "How do I check if AI crawlers can access my site?" Recall of a product's stated function vs. discovery of any tool that solves the problem "Does TrustGrowth check for AI visibility?" "Tools to measure a site's citation rate in ChatGPT and Claude" Confirmation of an existing association vs. an open-ended competitive surfaceUnits: n/a — illustrative prompt pairs, not measured results. Source: TrustGrowth prompt-design framework.
Why this split decides your score, not just describes it
An AI visibility score is a weighted average across a prompt set, so the branded/unbranded ratio determines whether the number measures recall or discovery before a single prompt is run. Hypothetically, if a prompt set is built from 80% branded queries, a well-known brand will post a high score almost by construction, because branded prompts mostly test whether the model has seen the brand name before. For a new product with limited public footprint, the same mix produces a score that looks usably high on branded prompts and collapses on the unbranded slice; an average that hides the real problem rather than surfacing it. (The distinction mirrors branded vs non-branded keywords: branded terms measure existing recognition, unbranded terms measure new discovery.)
Even a balanced 50/50 prompt set is a sample, not a census, and that limitation has to be disclosed in the same sentence as the score, not buried in a footnote. A set of 40 prompts run once on one date tells you about that day, those models, and that vocabulary, and nothing more. This is part of why AI visibility is treated as its own scored dimension rather than folded into a generic "SEO score"; see why GEO belongs in the TrustGrowth score for how classic search visibility and AI visibility are kept as separate constructs.
Branded prompts test recall, unbranded prompts test discovery
A branded prompt answers one question: does the model already know and vouch for us when asked directly. An unbranded prompt answers a different question: does the model surface us competitively when a stranger describes their problem in category language, with no hint of our name. Averaging the two into a single unlabeled number destroys the measurement's integrity: it's like averaging a thermometer reading with a barometer reading and calling it a "weather score." A tool that reports one blended AI visibility percentage without stating the branded/unbranded split is asking you to trust a number it hasn't shown you how it built.
How the two buckets are separated when measuring AI visibility
The method starts with two separately sourced prompt buckets, not one mixed list. Branded prompts are pulled from real branded queries already appearing in a site's Google Search Console query report, so the branded bucket reflects language actual users typed. Unbranded prompts are written separately as category-language questions modeled on how a buyer with no prior brand knowledge would phrase the same problem, matched to the site's actual product vocabulary rather than generic SEO boilerplate.
- Data source: GSC branded queries (branded bucket) plus hand-written category prompts (unbranded bucket)
- Models tested: Claude, GPT, and Gemini, each queried independently
- Reporting: results timestamped per run and reported per model — never averaged into one hidden cross-model blend
- Minimum sample: each bucket is sized and disclosed alongside the score
Full scoring mechanics live under AI visibility measurement across Claude, GPT, and Gemini, which reports recall as a rate — broken out by prompt intent and by engine, with a confidence interval — rather than a single pass/fail flag. When a bucket's sample is too small to separate signal from noise, so the interval comes out wider than the rate itself, the result is withheld as "sample too small to call" instead of being forced to a number. That distinction matters: a non-mention ("the model didn't surface us") is a countable result, while a too-thin sample is the absence of one, and collapsing the two into a single figure hides which you are looking at.
Common ways prompt design gets gamed
Four patterns inflate a score without improving what it measures:
- Overweighting branded prompts. A set skewed 70–80% branded (illustratively) posts an artificially high number for any brand with existing name recognition.
- Reporting one blended score. A single "AI visibility: 82%" headline (illustrative) with no disclosed ratio or sample size can't be audited or reproduced, unlike AI visibility measurement that reports recall per model with its sample and interval.
- Testing one model and generalizing. A score built on GPT alone says nothing about Claude or Gemini, which do return different answers to the identical prompt on the identical date.
- Forcing a too-thin sample to a number. Coding a bucket the sample can't support as a flat 0% or a pass reports certainty the data doesn't carry; a result too small to call belongs withheld, not rounded to zero.
This is the pattern covered in scores that quietly lie to you: a number without a disclosed method is a marketing claim wearing a measurement's clothing.
A framework for building an honest prompt set
- Pull real branded queries from GSC. Filter the query report for a manual brand-term list; these form the branded bucket.
- Write category-language prompts with no brand mention. Match wording to actual buyer vocabulary — support tickets, sales notes, onboarding survey text — rather than guessing.
- Set a minimum sample size per bucket and disclose it. A 5-prompt branded bucket and a 40-prompt unbranded bucket are not comparable; state both n values every time.
- Run identical prompts across every model in scope, on the same date. Report per-model results — Claude, GPT, Gemini — separately before any aggregate figure appears.
- Report each bucket's recall rate with its interval, and withhold what's too small to call. A bucket the sample can't support is information, not noise — mark it "sample too small to call" rather than scoring it a zero.
This is the discipline behind how we keep the score honest: disclose the method before the number, every time.
Limitations
Two limitations apply to any branded/unbranded prompt framework. First, prompt sets sample model behavior on a given date, and Claude, GPT, and Gemini are updated on schedules outside your control, so a recall rate measured today does not ensure the same rate next month. Second, category-language prompts are written by a human approximating buyer vocabulary, which introduces selection bias — two people writing the unbranded bucket for one product will not write identical prompts, and no process here eliminates that variance.
FAQ
What is the difference between a branded and unbranded AI prompt?
A branded prompt names the brand or product directly, testing whether a model already recalls and vouches for it. An unbranded prompt describes a problem in category language with no brand name, testing whether the model surfaces that brand competitively among unnamed alternatives.
Why does the branded/unbranded ratio affect an AI visibility score?
Because the score is a weighted average across the prompt set, a set skewed toward branded prompts produces a higher score for any brand with existing name recognition, regardless of how it performs when a buyer describes the problem without naming it.
Can a single AI visibility score be trusted without a disclosed ratio?
No. Without a disclosed branded/unbranded ratio, sample size, model list, and test date, a single blended percentage cannot be reproduced or audited.
Should AI visibility be tested on more than one model?
Yes. Claude, GPT, and Gemini can return different answers to an identical prompt on the same date, so a score built on one model cannot be generalized to "AI search" as a whole.
What happens when a prompt bucket returns too little data to score?
When a bucket's sample is too small to separate a real recall rate from noise — the confidence interval comes out wider than the rate — the honest handling is to withhold it as "sample too small to call," not convert it to a zero or a pass. A withheld bucket is information about sample size, not a measured failure.
Key takeaways
- A branded prompt names the brand directly; an unbranded prompt describes the problem in category language with no brand name.
- The branded/unbranded ratio in a prompt set determines whether a score measures recall or discovery, before any prompt is run.
- Even a balanced 50/50 prompt set is a sample, not a census — disclose sample size and test date alongside every score.
- Report recall as a rate with its confidence interval, and withhold any bucket too small to call rather than forcing it to a zero.
- Test identical prompts across every model in scope (Claude, GPT, Gemini) on the same date, and report results per model before any aggregate.
Know your site's real SEO score
Free GSC-verified audit, E-E-A-T scoring, and AI-powered content strategy.
Get Started Free