AI Search Visibility KPIs: 8 Metrics That Separate Mentions from Citations

AI search visibility is the measured presence of a brand or its content in a declared sample of AI-generated answers. A useful KPI set does not collapse that presence into one score. It records whether prompts produce brand mentions, whether responses cite the owned domain, how the brand compares with competitors, how it is described, which pages earn citations, and how often mentions and citations coincide. Every result is a sampled, directional estimate tied to specific prompts, engines, locations, dates, and repeated runs.

The word visibility hides two different outcomes. A mention occurs when generated prose names the brand. A citation occurs when the answer visibly presents a page or domain as a source. Either can happen without the other: an answer may recommend a brand while citing a review site, or it may cite the brand’s documentation without naming the company in the prose.

Neither outcome is a referral. A referral requires a person to click and reach the measured site. Nor is every page retrieved during answer generation a citation. Ahrefs’ measurement documentation explicitly separates pages that were visibly cited from pages that were found in the background but not cited.

A response-level mention, a visible source citation, a background retrieval, and a recorded referral are four different events. They require separate counting rules and separate data sources.

There is also no standardized AI visibility score. Ahrefs defines AI share of voice using estimated prompt impressions, while Semrush describes brand-performance share of voice as a percentage of mentions and exposes other visibility scores for other reports. Those measures can be useful inside their own systems, but the same label does not guarantee the same numerator, denominator, prompt corpus, weighting, or competitor set.

Use transparent response-level formulas instead. Start with a frozen roster of tracked prompts and a set of eligible responses collected from declared engines and repeated runs. Prompt mention coverage equals prompts that produce at least one brand mention across those runs divided by tracked prompts. Response mention rate equals eligible responses that mention the brand divided by all eligible responses. Mention share of voice equals the brand’s mention events divided by mention events for every brand in the fixed competitor set. Favorable mention rate equals favorable brand-mention responses divided by the brand-mention responses that were manually reviewed.

The citation side needs different denominators. Owned-domain citation rate equals eligible responses visibly citing the owned domain divided by all eligible responses. Citation share equals owned-domain citation events divided by citation events for all domains in the same response set, counting a domain at most once per response. Cited-page breadth is the number of distinct canonical owned URLs cited during the period. Mention-to-citation overlap equals responses that both mention the brand and cite its owned domain divided by responses that mention the brand.

Here is a clearly labeled illustrative calculation, not real company data. Suppose 40 frozen prompts are run on three engines twice, producing 240 eligible responses. The unnamed brand is mentioned for 24 prompts and in 72 responses: prompt mention coverage is 24 ÷ 40 = 60%, while response mention rate is 72 ÷ 240 = 30%. It accounts for 72 of 288 competitor-set brand mention events, so mention share of voice is 25%. Manual review marks 54 of its 72 mention-bearing responses favorable, making favorable mention rate 75%.

Its domain is visibly cited in 36 of 240 responses, so owned-domain citation rate is 15%. Those 36 domain events represent 36 of 360 total domain citation events, making citation share 10%. Nine distinct canonical owned URLs appear, so cited-page breadth is 9. Only 24 responses both mention the brand and cite its domain, so mention-to-citation overlap is 24 ÷ 72 = 33.3%. The gap is the point: 48 responses mention the brand without citing its site, while 12 cite its site without naming the brand.

InferredThese eight formulas are a transparent measurement specification derived from documented mention, citation, share-of-voice, sentiment, and fixed-prompt concepts. They are not presented as a cross-vendor standard.

The eight KPIs answer eight different questions

The value of this set is not the length of the dashboard. It is the discipline of keeping unlike events apart.

KPICalculationDecision it helps withWhat it cannot establish
1. Prompt mention coveragePrompts with at least one mention ÷ tracked promptsWhether the brand appears across the intended topic spaceWhether mentions recur reliably on repeated runs
2. Response mention rateBrand-mention responses ÷ eligible responsesHow consistently the brand appears in the sampled answersWhether its description is favorable or accurate
3. Mention share of voiceBrand mention events ÷ all fixed-set brand mention eventsWhether the brand is gaining or losing presence against declared competitorsMarket-wide awareness or a comparison with a different competitor set
4. Favorable mention rateFavorable reviewed mention responses ÷ manually reviewed mention responsesWhether answer narratives help or hurt the intended positioningObjective factual accuracy unless accuracy is coded separately
5. Owned-domain citation rateResponses citing the owned domain ÷ eligible responsesHow often the brand’s site is used as a visible sourceWhether the answer names the brand or sends a visitor
6. Citation shareOwned-domain citation events ÷ all domain citation eventsHow much of the visible source set the owned domain earnsPage importance, authority, placement, or traffic
7. Cited-page breadthDistinct canonical owned URLs citedWhether source visibility is distributed across useful assetsFrequency; one URL cited once and one cited repeatedly each count as one
8. Mention-to-citation overlapMention-and-owned-citation responses ÷ brand-mention responsesHow often brand visibility and owned evidence coincideCausality between the citation and the wording of the mention

1. Prompt mention coverage maps topical reach

Prompt mention coverage uses the prompt as its unit. It answers, “Across the questions we deliberately care about, where does the brand appear at all?” That makes it useful for locating category gaps: a team can segment the roster by problem, use case, comparison, role, or buying stage and see where coverage is absent.

Its limit is stability. If one of six runs mentions the brand for a prompt, that prompt is covered even though the appearance is fragile. Always pair coverage with response mention rate, and retain the number of runs behind the result.

2. Response mention rate measures repeatability

Response mention rate uses each eligible answer as its unit. It distinguishes a prompt that produced one isolated appearance from one that names the brand across engines and runs. Report it by engine, prompt segment, locale, and collection window before showing a blended total; an aggregate improvement can otherwise hide a loss on the surface that matters most.

This is still presence, not quality. A high response mention rate can coexist with unfavorable, outdated, or factually wrong descriptions.

3. Mention share of voice makes the competitor set explicit

A transparent mention-based share of voice asks what proportion of declared competitor-brand events belong to the brand. The competitor list must be frozen with the prompt roster. Adding weak competitors or removing a strong one can improve the KPI without changing a single answer.

Disclose whether one brand counts once per response, once per prompt, or every time its name appears. Response-level deduplication is usually easier to audit and less sensitive to repetitive prose. Do not compare this formula directly with Ahrefs’ impression-weighted version or any vendor score whose weighting is not identical.

4. Favorable mention rate adds narrative quality

Presence can be commercially unhelpful when the answer frames the brand for the wrong audience, repeats a retired claim, or describes a real limitation without context. Favorable mention rate adds a human-coded view of the narrative.

Write the rubric before reading the answers. Define favorable, neutral, unfavorable, mixed, and unscorable; record the evidence span; and review a sample twice to test whether two reviewers apply the categories consistently. Favorability is not the same as accuracy. If factual correctness matters, add a separate accuracy field rather than silently treating positive wording as true.

5. Owned-domain citation rate measures source inclusion

Owned-domain citation rate asks how often an eligible answer visibly uses the brand’s own site as a source. Count the domain once per response even if the answer cites several owned pages. That prevents a reference-heavy response from overwhelming the rest of the sample.

Bing’s AI Performance public-preview documentation is unusually clear about the boundary: its citation totals show displayed sources, not a page’s ranking, authority, importance, or placement. The same restraint belongs in an internal scorecard.

A domain-level citation event can be deduplicated within one response, and a citation count records source display rather than page importance or a user visit.

6. Citation share shows relative source presence

Citation rate uses responses as its denominator. Citation share instead uses all visibly cited domains in the same response set. It answers, “Of the source opportunities observed, what share did our domain receive?”

State whether “all domains” includes publishers, communities, regulators, documentation sites, and competitors. It usually should: restricting the denominator to commercial rivals changes the question from overall source presence to competitive-domain presence. Either version can be valid, but they need different names and cannot share a trend line.

7. Cited-page breadth exposes concentration risk

Cited-page breadth is a count, not a percentage. Canonicalize URLs before deduplication so tracking parameters, fragments, protocol variants, and alternate hostnames do not inflate it. Report breadth beside the citation distribution: ten cited pages look healthy until one page accounts for nearly every citation.

Breadth helps content owners see whether AI source visibility reaches product documentation, definitions, research, comparison material, and other intended evidence assets. It does not imply that more cited URLs are always better. A small, authoritative library can be the right result for a narrow prompt set.

8. Mention-to-citation overlap connects the two systems

Overlap is the metric that prevents a team from treating mentions and citations as synonyms. Build a response-level two-by-two table before calculating it:

Owned domain citedOwned domain not cited
Brand mentionedBrand presence and owned evidence coincideBrand is visible, but the owned site is not a displayed source
Brand not mentionedOwned content supports the answer without an explicit brand mentionNeither observed in that response

The overlap formula uses brand-mention responses as its denominator, so it answers, “When we are named, how often is our own evidence also visible?” A different but legitimate question—“When we are cited, how often are we named?”—reverses the denominator. Label it separately if you need it.

Never put a mention numerator over an unlabeled citation denominator. The denominator is the metric’s meaning.

Instrument the sample before interpreting a trend

A defensible KPI starts with a measurement contract. Record these fields before the first collection:

  • Prompt roster and version: the exact prompts, topic labels, intended audience context, and inclusion rationale. Freeze the roster for period comparisons; Semrush likewise warns that expanding a prompt set mid-cycle can inflate mentions without proving improvement.
  • Surfaces and conditions: engine, model or product surface, locale, location when controllable, account or personalization state, collection time, and run count.
  • Eligibility rules: what qualifies as a completed answer, how refusals and errors are treated, and whether reruns replace or add observations.
  • Entity rules: official name, aliases, products, ambiguous terms, parent brands, and the response-level deduplication rule.
  • Citation rules: what counts as visibly cited, whether redirected URLs resolve to a canonical domain, and how domain and page events are deduplicated.
  • Competitor set: the fixed entities included in share of voice, with a change log for additions, removals, and mergers.
  • Review rubric: sentiment labels, factual-accuracy fields, unscorable cases, reviewer identity, and adjudication method.
  • Evidence receipt: raw response, visible sources, timestamp, engine, prompt ID, run ID, parser version, and any manual correction.

This contract matters because generative answers are not deterministic. A June 2026 repeated-sampling preprint collected results across three generative-search platforms and three consumer-product topics. It found substantial citation variability and showed that many apparent domain differences fell within the measurement noise.

The cited study found that identical queries can return different cited sources and that single-run citation shares can look more precise and stable than repeated samples justify.

That study does not create a universal minimum run count. It does justify three reporting habits: repeat observations, show the numerator and denominator with every rate, and present an uncertainty interval or repeated-run range when the sample supports one. Flag small samples instead of decorating them with extra decimal places. A movement from one period to another is actionable only when the collection contract stayed fixed and the change is larger than ordinary run-to-run variation.

Keep sampled visibility, first-party impressions, and referrals separate

No single data source observes the whole journey. By the June 27, 2026 evidence cutoff, platform owners had begun exposing useful first-party views, but each view measured a different surface.

Google announced dedicated generative AI performance reports in Search Console on June 3, initially for a subset of websites. The announced reports show impressions and pages for AI Overviews, AI Mode, and generative AI features in Discover, with country, date, and—on Search—device dimensions. This is first-party Google visibility data. It is not the same sample as a cross-engine prompt tracker and does not count brand mentions in generated prose.

Bing’s AI Performance report covers supported Microsoft AI experiences and selected partner integrations. It reports total citations, average cited pages, sampled grounding queries, page-level citation activity, and trends. These first-party totals can validate where an owned site is being cited within Bing’s declared scope, but they should retain their own denominator and trend line.

Google Analytics added an AI Assistant channel on May 13. Recognized assistant referrers receive the ai-assistant medium and (ai-assistant) campaign. This is the right layer for recorded visits, engaged sessions, key events, and downstream conversion analysis. It cannot observe people who saw a mention but did not click, arrived later without a referrer, or used an assistant the classification did not recognize.

Evidence layerWhat it observesKeep it forDo not infer
Fixed-prompt trackerSampled answer text and displayed sources across declared runsThe eight visibility KPIs in this articleComplete user demand or platform-wide exposure
Google Search Console generative AI reportFirst-party impressions and pages on declared Google AI surfacesGoogle-specific owned-page visibilityCross-engine mentions, citations, or referrals
Bing AI PerformanceFirst-party citations and cited pages on supported Microsoft surfacesMicrosoft-specific source visibilityRanking, authority, placement, clicks, or market-wide share
Web analytics AI channelRecognized referral visits and on-site behaviorSessions, key events, and conversions after a recorded clickAnswer exposure without a visit or causal credit for every later outcome
First-party search reports and web analytics add valuable evidence, but their scopes and observed events differ from one another and from a fixed cross-engine response sample.

Read the gaps, not just the green arrows

The most useful finding is often the disagreement between metrics.

PatternPlausible readingNext check
Prompt coverage rises; response mention rate does notThe brand entered more topics, but appearances may be sporadicSegment new coverage by prompt and inspect repeated-run stability
Mentions rise; owned-domain citation rate stays flatThird-party sources or uncited brand knowledge may be driving visibilityInspect the cited domains and the mention-without-owned-citation cells
Owned citations rise; mentions stay flatThe content may support answers without transferring explicit brand presenceReview cited passages, page templates, and citation-without-mention responses
Mention share of voice rises; absolute coverage fallsThe whole competitor set may have lost presence while the brand lost lessShow event counts and prompt coverage beside the share
Citation rate rises; cited-page breadth fallsSource visibility may be concentrating on fewer pagesPlot citations by canonical URL and test dependency on the leading page
Mention-to-citation overlap rises; favorability fallsThe brand and its evidence co-occur more often, but the narrative may be worseningReview sentiment evidence spans and factual accuracy separately

These are hypotheses, not automatic diagnoses. The raw answer and source receipt should remain one click away from every aggregate so an operator can see what actually changed.

There is no universal “good” AI visibility score

No broadly accepted benchmark applies across industries, prompt rosters, engines, countries, or tools. The vendor definitions alone prevent a clean comparison: Ahrefs’ AI share of voice uses estimated impressions, while Semrush documents a mention-based brand share of voice and separate visibility scores. Repeated responses add another source of variation.

A useful benchmark is therefore internal and decision-bound. Freeze the measurement contract, establish a baseline, compare the same segments over time, and keep a stable competitor set. Call a result good only in context: priority-prompt coverage is broad enough for the decision, response-level presence is repeatable, narrative quality is acceptable, and citation gains persist beyond ordinary sampling variation. Publish the observed counts and range so another person can challenge that judgment.

Build the smallest scorecard that survives review

Group the eight KPIs into four lines rather than blending them:

  1. Brand presence: prompt mention coverage and response mention rate.
  2. Competitive narrative: mention share of voice and favorable mention rate.
  3. Owned evidence: owned-domain citation rate, citation share, and cited-page breadth.
  4. Connection: mention-to-citation overlap.

For each KPI, show the current numerator and denominator, prior-period result, absolute change, repeated-run range or uncertainty interval, and the segment responsible for the movement. Put prompt-set version, engines, locales, run count, competitor-set version, and collection dates in the scorecard header. Place referral sessions and conversions in a separate downstream panel.

The executive headline should be a finding, not a blended score: “Brand mentions broadened across priority prompts, but owned citations remain concentrated on one page” is more actionable than “AI visibility rose six points.” It tells content, digital PR, analytics, and leadership which part of the system moved and which part did not.

The decision
Use these eight KPIs when the decision concerns presence inside AI-generated answers.

Use first-party platform reports to validate their own surfaces, and use analytics to measure the visits and outcomes that follow recorded clicks. If the team cannot state the prompt set, engine mix, counting rule, and denominator, do not optimize the number yet; fix the measurement contract first.

Continue the evidence path

Run your growth team from one screen.

Invite only