AI Search Visibility KPIs: 8 Metrics That Separate Mentions from Citations
AI search visibility is the measured presence of a brand or its content in a declared sample of AI-generated answers. A useful KPI set does not collapse that presence into one score. It records whether prompts produce brand mentions, whether responses cite the owned domain, how the brand compares with competitors, how it is described, which pages earn citations, and how often mentions and citations coincide. Every result is a sampled, directional estimate tied to specific prompts, engines, locations, dates, and repeated runs.
The word visibility hides two different outcomes. A mention occurs when generated prose names the brand. A citation occurs when the answer visibly presents a page or domain as a source. Either can happen without the other: an answer may recommend a brand while citing a review site, or it may cite the brand’s documentation without naming the company in the prose.
Neither outcome is a referral. A referral requires a person to click and reach the measured site. Nor is every page retrieved during answer generation a citation. Ahrefs’ measurement documentation explicitly separates pages that were visibly cited from pages that were found in the background but not cited.
There is also no standardized AI visibility score. Ahrefs defines AI share of voice using estimated prompt impressions, while Semrush describes brand-performance share of voice as a percentage of mentions and exposes other visibility scores for other reports. Those measures can be useful inside their own systems, but the same label does not guarantee the same numerator, denominator, prompt corpus, weighting, or competitor set.
Use transparent response-level formulas instead. Start with a frozen roster of tracked prompts and a set of eligible responses collected from declared engines and repeated runs. Prompt mention coverage equals prompts that produce at least one brand mention across those runs divided by tracked prompts. Response mention rate equals eligible responses that mention the brand divided by all eligible responses. Mention share of voice equals the brand’s mention events divided by mention events for every brand in the fixed competitor set. Favorable mention rate equals favorable brand-mention responses divided by the brand-mention responses that were manually reviewed.
The citation side needs different denominators. Owned-domain citation rate equals eligible responses visibly citing the owned domain divided by all eligible responses. Citation share equals owned-domain citation events divided by citation events for all domains in the same response set, counting a domain at most once per response. Cited-page breadth is the number of distinct canonical owned URLs cited during the period. Mention-to-citation overlap equals responses that both mention the brand and cite its owned domain divided by responses that mention the brand.
Here is a clearly labeled illustrative calculation, not real company data. Suppose 40 frozen prompts are run on three engines twice, producing 240 eligible responses. The unnamed brand is mentioned for 24 prompts and in 72 responses: prompt mention coverage is 24 ÷ 40 = 60%, while response mention rate is 72 ÷ 240 = 30%. It accounts for 72 of 288 competitor-set brand mention events, so mention share of voice is 25%. Manual review marks 54 of its 72 mention-bearing responses favorable, making favorable mention rate 75%.
Its domain is visibly cited in 36 of 240 responses, so owned-domain citation rate is 15%. Those 36 domain events represent 36 of 360 total domain citation events, making citation share 10%. Nine distinct canonical owned URLs appear, so cited-page breadth is 9. Only 24 responses both mention the brand and cite its domain, so mention-to-citation overlap is 24 ÷ 72 = 33.3%. The gap is the point: 48 responses mention the brand without citing its site, while 12 cite its site without naming the brand.
The eight KPIs answer eight different questions
The value of this set is not the length of the dashboard. It is the discipline of keeping unlike events apart.
| KPI | Calculation | Decision it helps with | What it cannot establish |
|---|---|---|---|
| 1. Prompt mention coverage | Prompts with at least one mention ÷ tracked prompts | Whether the brand appears across the intended topic space | Whether mentions recur reliably on repeated runs |
| 2. Response mention rate | Brand-mention responses ÷ eligible responses | How consistently the brand appears in the sampled answers | Whether its description is favorable or accurate |
| 3. Mention share of voice | Brand mention events ÷ all fixed-set brand mention events | Whether the brand is gaining or losing presence against declared competitors | Market-wide awareness or a comparison with a different competitor set |
| 4. Favorable mention rate | Favorable reviewed mention responses ÷ manually reviewed mention responses | Whether answer narratives help or hurt the intended positioning | Objective factual accuracy unless accuracy is coded separately |
| 5. Owned-domain citation rate | Responses citing the owned domain ÷ eligible responses | How often the brand’s site is used as a visible source | Whether the answer names the brand or sends a visitor |
| 6. Citation share | Owned-domain citation events ÷ all domain citation events | How much of the visible source set the owned domain earns | Page importance, authority, placement, or traffic |
| 7. Cited-page breadth | Distinct canonical owned URLs cited | Whether source visibility is distributed across useful assets | Frequency; one URL cited once and one cited repeatedly each count as one |
| 8. Mention-to-citation overlap | Mention-and-owned-citation responses ÷ brand-mention responses | How often brand visibility and owned evidence coincide | Causality between the citation and the wording of the mention |
1. Prompt mention coverage maps topical reach
Prompt mention coverage uses the prompt as its unit. It answers, “Across the questions we deliberately care about, where does the brand appear at all?” That makes it useful for locating category gaps: a team can segment the roster by problem, use case, comparison, role, or buying stage and see where coverage is absent.
Its limit is stability. If one of six runs mentions the brand for a prompt, that prompt is covered even though the appearance is fragile. Always pair coverage with response mention rate, and retain the number of runs behind the result.
2. Response mention rate measures repeatability
Response mention rate uses each eligible answer as its unit. It distinguishes a prompt that produced one isolated appearance from one that names the brand across engines and runs. Report it by engine, prompt segment, locale, and collection window before showing a blended total; an aggregate improvement can otherwise hide a loss on the surface that matters most.
This is still presence, not quality. A high response mention rate can coexist with unfavorable, outdated, or factually wrong descriptions.
3. Mention share of voice makes the competitor set explicit
A transparent mention-based share of voice asks what proportion of declared competitor-brand events belong to the brand. The competitor list must be frozen with the prompt roster. Adding weak competitors or removing a strong one can improve the KPI without changing a single answer.
Disclose whether one brand counts once per response, once per prompt, or every time its name appears. Response-level deduplication is usually easier to audit and less sensitive to repetitive prose. Do not compare this formula directly with Ahrefs’ impression-weighted version or any vendor score whose weighting is not identical.
4. Favorable mention rate adds narrative quality
Presence can be commercially unhelpful when the answer frames the brand for the wrong audience, repeats a retired claim, or describes a real limitation without context. Favorable mention rate adds a human-coded view of the narrative.
Write the rubric before reading the answers. Define favorable, neutral, unfavorable, mixed, and unscorable; record the evidence span; and review a sample twice to test whether two reviewers apply the categories consistently. Favorability is not the same as accuracy. If factual correctness matters, add a separate accuracy field rather than silently treating positive wording as true.
5. Owned-domain citation rate measures source inclusion
Owned-domain citation rate asks how often an eligible answer visibly uses the brand’s own site as a source. Count the domain once per response even if the answer cites several owned pages. That prevents a reference-heavy response from overwhelming the rest of the sample.
Bing’s AI Performance public-preview documentation is unusually clear about the boundary: its citation totals show displayed sources, not a page’s ranking, authority, importance, or placement. The same restraint belongs in an internal scorecard.
6. Citation share shows relative source presence
Citation rate uses responses as its denominator. Citation share instead uses all visibly cited domains in the same response set. It answers, “Of the source opportunities observed, what share did our domain receive?”
State whether “all domains” includes publishers, communities, regulators, documentation sites, and competitors. It usually should: restricting the denominator to commercial rivals changes the question from overall source presence to competitive-domain presence. Either version can be valid, but they need different names and cannot share a trend line.
7. Cited-page breadth exposes concentration risk
Cited-page breadth is a count, not a percentage. Canonicalize URLs before deduplication so tracking parameters, fragments, protocol variants, and alternate hostnames do not inflate it. Report breadth beside the citation distribution: ten cited pages look healthy until one page accounts for nearly every citation.
Breadth helps content owners see whether AI source visibility reaches product documentation, definitions, research, comparison material, and other intended evidence assets. It does not imply that more cited URLs are always better. A small, authoritative library can be the right result for a narrow prompt set.
8. Mention-to-citation overlap connects the two systems
Overlap is the metric that prevents a team from treating mentions and citations as synonyms. Build a response-level two-by-two table before calculating it:
| Owned domain cited | Owned domain not cited | |
|---|---|---|
| Brand mentioned | Brand presence and owned evidence coincide | Brand is visible, but the owned site is not a displayed source |
| Brand not mentioned | Owned content supports the answer without an explicit brand mention | Neither observed in that response |
The overlap formula uses brand-mention responses as its denominator, so it answers, “When we are named, how often is our own evidence also visible?” A different but legitimate question—“When we are cited, how often are we named?”—reverses the denominator. Label it separately if you need it.
Instrument the sample before interpreting a trend
A defensible KPI starts with a measurement contract. Record these fields before the first collection:
- Prompt roster and version: the exact prompts, topic labels, intended audience context, and inclusion rationale. Freeze the roster for period comparisons; Semrush likewise warns that expanding a prompt set mid-cycle can inflate mentions without proving improvement.
- Surfaces and conditions: engine, model or product surface, locale, location when controllable, account or personalization state, collection time, and run count.
- Eligibility rules: what qualifies as a completed answer, how refusals and errors are treated, and whether reruns replace or add observations.
- Entity rules: official name, aliases, products, ambiguous terms, parent brands, and the response-level deduplication rule.
- Citation rules: what counts as visibly cited, whether redirected URLs resolve to a canonical domain, and how domain and page events are deduplicated.
- Competitor set: the fixed entities included in share of voice, with a change log for additions, removals, and mergers.
- Review rubric: sentiment labels, factual-accuracy fields, unscorable cases, reviewer identity, and adjudication method.
- Evidence receipt: raw response, visible sources, timestamp, engine, prompt ID, run ID, parser version, and any manual correction.
This contract matters because generative answers are not deterministic. A June 2026 repeated-sampling preprint collected results across three generative-search platforms and three consumer-product topics. It found substantial citation variability and showed that many apparent domain differences fell within the measurement noise.
That study does not create a universal minimum run count. It does justify three reporting habits: repeat observations, show the numerator and denominator with every rate, and present an uncertainty interval or repeated-run range when the sample supports one. Flag small samples instead of decorating them with extra decimal places. A movement from one period to another is actionable only when the collection contract stayed fixed and the change is larger than ordinary run-to-run variation.
Keep sampled visibility, first-party impressions, and referrals separate
No single data source observes the whole journey. By the June 27, 2026 evidence cutoff, platform owners had begun exposing useful first-party views, but each view measured a different surface.
Google announced dedicated generative AI performance reports in Search Console on June 3, initially for a subset of websites. The announced reports show impressions and pages for AI Overviews, AI Mode, and generative AI features in Discover, with country, date, and—on Search—device dimensions. This is first-party Google visibility data. It is not the same sample as a cross-engine prompt tracker and does not count brand mentions in generated prose.
Bing’s AI Performance report covers supported Microsoft AI experiences and selected partner integrations. It reports total citations, average cited pages, sampled grounding queries, page-level citation activity, and trends. These first-party totals can validate where an owned site is being cited within Bing’s declared scope, but they should retain their own denominator and trend line.
Google Analytics added an AI Assistant channel on May 13. Recognized assistant referrers receive the ai-assistant medium and (ai-assistant) campaign. This is the right layer for recorded visits, engaged sessions, key events, and downstream conversion analysis. It cannot observe people who saw a mention but did not click, arrived later without a referrer, or used an assistant the classification did not recognize.
| Evidence layer | What it observes | Keep it for | Do not infer |
|---|---|---|---|
| Fixed-prompt tracker | Sampled answer text and displayed sources across declared runs | The eight visibility KPIs in this article | Complete user demand or platform-wide exposure |
| Google Search Console generative AI report | First-party impressions and pages on declared Google AI surfaces | Google-specific owned-page visibility | Cross-engine mentions, citations, or referrals |
| Bing AI Performance | First-party citations and cited pages on supported Microsoft surfaces | Microsoft-specific source visibility | Ranking, authority, placement, clicks, or market-wide share |
| Web analytics AI channel | Recognized referral visits and on-site behavior | Sessions, key events, and conversions after a recorded click | Answer exposure without a visit or causal credit for every later outcome |
Read the gaps, not just the green arrows
The most useful finding is often the disagreement between metrics.
| Pattern | Plausible reading | Next check |
|---|---|---|
| Prompt coverage rises; response mention rate does not | The brand entered more topics, but appearances may be sporadic | Segment new coverage by prompt and inspect repeated-run stability |
| Mentions rise; owned-domain citation rate stays flat | Third-party sources or uncited brand knowledge may be driving visibility | Inspect the cited domains and the mention-without-owned-citation cells |
| Owned citations rise; mentions stay flat | The content may support answers without transferring explicit brand presence | Review cited passages, page templates, and citation-without-mention responses |
| Mention share of voice rises; absolute coverage falls | The whole competitor set may have lost presence while the brand lost less | Show event counts and prompt coverage beside the share |
| Citation rate rises; cited-page breadth falls | Source visibility may be concentrating on fewer pages | Plot citations by canonical URL and test dependency on the leading page |
| Mention-to-citation overlap rises; favorability falls | The brand and its evidence co-occur more often, but the narrative may be worsening | Review sentiment evidence spans and factual accuracy separately |
These are hypotheses, not automatic diagnoses. The raw answer and source receipt should remain one click away from every aggregate so an operator can see what actually changed.
There is no universal “good” AI visibility score
No broadly accepted benchmark applies across industries, prompt rosters, engines, countries, or tools. The vendor definitions alone prevent a clean comparison: Ahrefs’ AI share of voice uses estimated impressions, while Semrush documents a mention-based brand share of voice and separate visibility scores. Repeated responses add another source of variation.
A useful benchmark is therefore internal and decision-bound. Freeze the measurement contract, establish a baseline, compare the same segments over time, and keep a stable competitor set. Call a result good only in context: priority-prompt coverage is broad enough for the decision, response-level presence is repeatable, narrative quality is acceptable, and citation gains persist beyond ordinary sampling variation. Publish the observed counts and range so another person can challenge that judgment.
Build the smallest scorecard that survives review
Group the eight KPIs into four lines rather than blending them:
- Brand presence: prompt mention coverage and response mention rate.
- Competitive narrative: mention share of voice and favorable mention rate.
- Owned evidence: owned-domain citation rate, citation share, and cited-page breadth.
- Connection: mention-to-citation overlap.
For each KPI, show the current numerator and denominator, prior-period result, absolute change, repeated-run range or uncertainty interval, and the segment responsible for the movement. Put prompt-set version, engines, locales, run count, competitor-set version, and collection dates in the scorecard header. Place referral sessions and conversions in a separate downstream panel.
The executive headline should be a finding, not a blended score: “Brand mentions broadened across priority prompts, but owned citations remain concentrated on one page” is more actionable than “AI visibility rose six points.” It tells content, digital PR, analytics, and leadership which part of the system moved and which part did not.
Use first-party platform reports to validate their own surfaces, and use analytics to measure the visits and outcomes that follow recorded clicks. If the team cannot state the prompt set, engine mix, counting rule, and denominator, do not optimize the number yet; fix the measurement contract first.
Continue the evidence path
Related reading
Read first
Generative Engine Optimization Terms: GEO, AEO, AI Search, and Citation Defined
Separate mentions, citations, answers, and rankings before assigning metrics to AI-search visibility work.
Related
GEO vs. SEO: Key Differences in Rankings, Mentions, and Citations
Compare the outcomes produced by traditional search and generative systems before combining their scorecards.
Next step
SEO for AI Search: What a Lean Team Can Reuse Before Adding New Work
Connect visibility measures to a lean operating plan that reuses sound SEO work before adding new activity.