LLM Visibility: How to Measure and Improve It
LLM visibility is the presence of a brand or its content in answers generated by large language models. The main things to measure are whether the answer names the brand and whether it cites the brand’s pages. Track those outcomes across a defined set of questions and AI experiences, then check what the answers actually say. Retrieval, visible citations, and visits to a website belong in separate measures.

What counts as LLM visibility?
For answer monitoring, distinguish four observations:
| Observation | Evidence to record | What it measures |
|---|---|---|
| Brand mention | The brand name in the answer text | Presence in the generated response |
| Retrieval | A page shown in an exposed retrieval or source record | A page encountered during answer preparation |
| Visible citation | A linked page presented as a source | Content displayed as supporting material |
| Referral visit | A session attributed to an AI assistant in site analytics | Traffic that reached the website |
These distinctions are reflected in Ahrefs’ Brand Radar documentation: a brand mention counts once per response, while retrieved pages and visibly cited pages are recorded separately. Its “Found in” category includes both cited and retrieved pages, so that label should not be read as a count of visible citations alone.
Measure brand presence and owned-site citations independently. For citations, also separate pages on the monitored domain from third-party pages that discuss the brand. This keeps the answer’s wording, the publisher of its evidence, and traffic attribution identifiable in the report.
LLM visibility is an outcome to measure. Improving it requires identifying the questions, answers, and source pages involved. A monitoring score alone does not explain which page needs attention or which claim needs correction.
How to measure mentions, citations, and prominence
Use explicit counting rules before collecting answers. The following definitions provide a reproducible basis for an audit; they are not a specification that every monitoring product follows.
| Metric | Counting rule for the audit |
|---|---|
| Brand mention rate | Valid responses naming the brand ÷ all valid responses × 100 |
| Owned-domain citation rate | Valid responses citing at least one page on the monitored domain ÷ all valid responses × 100 |
| Share of tracked brand mentions | The brand’s response-level mentions ÷ response-level mentions summed across the fixed set of tracked brands × 100 |
| Cited-page coverage | Number of distinct owned URLs visibly cited during the reporting period |
| Third-party source coverage | Distinct outside URLs and domains cited alongside claims about the brand |
Count each brand once per response, regardless of repeated names. For page-level citation frequency, count each URL once per response. For domain-level frequency, count the domain once even when several of its pages are cited. Preserve the underlying links so both views can be recalculated.
Publish numerators and denominators beside percentages. Define a valid response, record failed runs separately, and specify how domain aliases, subdomains, redirects, and brand-name variants are handled. Leave a percentage undefined when its denominator is zero.
Share of tracked brand mentions answers a different question from mention rate: its denominator is all recorded mentions of the selected brands. Multiple brands can appear in the same answer. Keep the tracked brand list fixed for comparisons over time.
Vendor labels need their own definitions. Bing’s Citation Share explanation defines a site’s share of citations across all sites for a specific grounding query; it does not represent traffic share or a ranking. Ahrefs’ AI share of voice calculation instead uses its impression estimates, which are derived from Google search volumes. Those estimates should not be relabeled as observed AI audience counts.
Record prominence separately from frequency. Useful fields include the first paragraph containing the brand, whether it appears in a heading, and its position in an explicitly ordered recommendation list. Preserve the surrounding text and label the mention’s role: recommendation, example, comparison, or warning. Use “not applicable” for list position in an answer without an ordered list. These fields describe presentation within an observed answer, rather than a universal LLM rank.
Build a repeatable question set
Start with questions relevant to the subject being monitored. Use existing search-query records, support questions, product documentation, and comparison topics as inputs. Group the questions by purpose, such as learning a definition, comparing approaches, checking suitability, or resolving a problem.
Separate branded questions, which contain the brand or product name, from non-branded questions, which ask about a category or problem without naming it. Report their results independently. The branded group tests descriptions of a named entity; the non-branded group tests whether that entity appears without being supplied in the question.
For each run, save:
- The exact prompt, topic group, and branded or non-branded classification.
- The platform and experience, including search or browsing mode and model label where exposed.
- The language, location settings, date, time, and conversation context.
- The full answer, counted brand names, citation URLs, and mention context.
- Exposed retrieval records, with unavailable retrieval data marked as unknown.
- Run status, including failures or responses excluded under the declared rules.
Use a consistent conversation setup for comparisons. Keep new-session tests separate from follow-up conversations. Version the prompt list and matching rules; adding questions or changing name matching creates a new measurement set.
Repeat prompts rather than treating one answer as a stable result. A 2026 preprint on AI visibility uncertainty, studying three platforms and three consumer-product topics, found substantial variation in cited sources across repeated samples. Its findings support reporting uncertainty, but do not establish one monitoring frequency or run count for every subject.
Choose a schedule and repeat count before collection. Report the number of valid runs per prompt, show variation across collection windows, and compare periods using the same prompt mix and conditions. Keep results separated by platform before presenting an aggregate. Any weighting across platforms or topics should be documented so a change in weighting cannot be mistaken for a change in visibility.
Which reports can track LLM visibility?
Combine answer-level records with reporting from the platforms being monitored. Each source measures a different part of visibility.
Brand and answer monitoring. Ahrefs’ AI visibility metrics include brand mentions, page and domain citations, and retrieval-related records. Inspect the actual responses alongside the summary metrics, and verify the monitored prompt corpus and entity settings before comparing totals.
Microsoft citation reporting. Bing Webmaster Tools’ AI Performance announcement documents citations, cited URLs, and grounding-query phrases across Microsoft Copilot, Bing AI summaries, and selected partner integrations. It states that page citation counts measure frequency rather than importance, ranking, or placement. Use the report to identify cited pages within that coverage.
Citations and referral traffic. Microsoft Clarity’s Citations Dashboard documents page citations, grounding queries, cited pages, and AI referral traffic. Its referral metric is the share of site sessions attributed to AI assistants. The documentation also notes that grounding queries can differ from users’ exact wording; keep those retrieval phrases distinct from the original prompts in an audit.
Google AI search impressions. Google’s Generative AI performance report for Search documents site-link impressions in AI Overviews and AI Mode, with page, country, device, and date dimensions. It excludes Search Labs experiments. These are impressions of links to the site in supported Google features, rather than counts of brand names in all LLM answers.
Manual review can supply the answer text and visible links for a declared prompt set. Platform reports add observations within their documented coverage. Keep these datasets separate unless their units, time periods, and coverage match. Do not add mentions, citations, link impressions, and visits into one total: they count different events.
Check the accuracy behind the visibility
Frequency measures presence. Check the associated statements to evaluate whether that presence is useful and accurate.
The study Evaluating Verifiability in Generative Search Engines distinguishes citation completeness from citation correctness: whether claims are supported by citations, and whether each citation supports its associated statement. The study evaluated four systems in 2023; its historical results should not be used as current accuracy rates.
Audit the answers and source pages directly:
- Identify the exact statement about the brand or topic.
- Open the citation attached to that statement.
- Locate the passage that supports, qualifies, or contradicts it.
- Record whether the answer preserves the source’s conditions and limitations.
- Mark unsupported claims and incorrect brand identification for correction.
Review uncited brand mentions against authoritative information as well. Keep the answer text, source passage, URL, and review date together. This makes the distinction between a missing mention, a misleading description, and an unsupported citation visible in the report.
Source coverage should include the publisher behind each cited page. Group owned documentation, independent editorial sources, directories, forums, and other source types separately. Use those groups to locate inaccurate information and gaps in the material being cited, without assigning an unsupported authority score to a domain.
How to improve LLM visibility
Use the measurement record to identify relevant pages, verify access, and correct information. Treat content changes as actions to evaluate with later observations, rather than guaranteed citation tactics.
Check eligibility and access
Verify the intended page’s crawlability, indexing, and content visibility against the documentation for the target experience. For Google, the generative AI optimization guide specifies indexing and snippet eligibility and retains foundational SEO practices. Eligibility does not guarantee that a page will appear.
Also check Google’s Search generative AI control in Search Console under Settings > Search generative AI. It manages inclusion in supported Google AI features. Check the effective setting, including inheritance from a parent property, before investigating missing Google AI visibility.
Make the relevant page a better source
Review the page against the question being monitored. Put the direct answer where it is easy to find, define names consistently, state important conditions beside the claims they qualify, and link factual assertions to verifiable evidence. Correct outdated information on owned pages and record inaccurate statements found on third-party sources separately.
Google’s content guidance for generative AI search emphasizes useful, distinctive content organized for readers. Apply that guidance to the identified information gap: improve the explanation, supporting evidence, or missing limitation on the relevant page. Record the page revision and date so subsequent observations can be compared with the baseline.
Recheck mentions, citations, and visits separately
Run the same questions again under the recorded conditions. Review both frequency and accuracy. Keep referral visits and site outcomes in their own analytics view; a citation count does not supply evidence that a visit occurred.
Report observed changes with the prompt set, platforms, collection dates, and valid run counts. Describe a change following an edit as an observation. Establishing that the edit caused the change requires evidence beyond a before-and-after count.
Frequently asked questions
Is a brand mention the same as a citation?
A mention names the brand in the answer; a citation links to a source page. Ahrefs’ counting definitions distinguish the two. Track mentions of the brand, citations of owned pages, and citations of third-party pages separately.
What is a good LLM visibility score?
Assess results against a declared question set and a comparable baseline. Include mention rate, citation rate, and answer accuracy, with the underlying counts. Compare vendor scores only after checking their definitions; Bing’s Citation Share guidance explicitly distinguishes its metric from ranking, traffic share, and quality scores.
How often should LLM visibility be checked?
Use a consistent schedule with repeated runs, and increase observation where uncertainty could affect a consequential decision. The AI visibility uncertainty preprint supports repeated sampling rather than reliance on a single response. State the schedule and sample size with the results instead of presenting them as universal requirements.
Does llms.txt improve LLM visibility?
Google’s official generative AI search guide states that Google Search ignores llms.txt and does not require special AI schema markup. That guidance applies to Google Search. Evaluate requirements for other systems using their own documentation.