LLM (large language model) visibility is the measure of how often, how prominently and how accurately models such as ChatGPT, Gemini, Claude, Perplexity, Copilot and Google AI Overviews mention, recommend and cite a brand when buyers ask about its category. It is not a ranking. Each answer is generated fresh, so the same prompt can name you on Monday and skip you on Tuesday. Measuring LLM visibility therefore means sampling many prompts, many times, across several engines, and reporting rates rather than positions.
Buyers now form shortlists inside answers you do not control. This guide is current as of October 2026. It defines the metrics, gives a method you can run without a tool, explains why single checks mislead and compares the main ways to measure, from spreadsheets to commercial platforms.
Key takeaways
- LLM visibility is measured as rates across repeated samples, such as mention rate, citation share and answer share, because AI answers change between runs.
- A credible measurement uses a fixed prompt library split by funnel stage, branded and non-branded intent, and engine, with each prompt run several times.
- Citation tracking misses part of the picture, since models often answer from training data without citing anything or cite an intermediary page instead of the original source.
- A visibility drop on one engine usually points to a retrieval or model change, while a drop across every engine usually points to a content or reputation gap.
- Monitoring shows where a brand is missing, but the score only moves when someone publishes the pages, earns the third-party mentions and fixes the sources the models read.
What LLM visibility actually measures
In SEO, a page holds a ranked position for a query, and position 3 in search today is usually position 3 tomorrow. LLM visibility measures your presence inside generated answers instead. In an AI answer, your brand can be recommended first, mentioned with a caveat, cited as a source without being named, or left out, and the outcome varies by run, engine and region. Each metric below is a ratio over a sample of answers, because that is the only defensible way to report a non-deterministic system.
| Metric | What it answers | Calculation logic |
|---|---|---|
| Mention rate (inclusion rate) | How often are we named at all? | Answers mentioning the brand ÷ total answers sampled |
| Citation rate | How often is our site used as a source? | Answers citing at least one owned URL ÷ answers that show citations |
| Citation share | What portion of the sources are ours? | Owned-domain citations ÷ all citations in the sample |
| Answer share (AI share of voice) | How much of the competitive conversation do we hold? | Brand mentions ÷ mentions of all tracked brands |
| Answer position | When named, are we first or an afterthought? | Mean rank of first mention among listed brands, only in answers that include us |
| Sentiment | How are we framed? | (Positive − negative mentions) ÷ total mentions, from −1 to +1 |
| Competitive displacement | Where does a rival take our place? | Prompts where a competitor appears and we do not ÷ total prompts |
| Prompt coverage | How much of the category do we reliably appear in? | Prompts where we appear in at least 3 of 5 runs ÷ total prompts |
Most teams start with mention rate as the headline number, and our guide on how often AI answers mention your brand covers it in depth. On its own, it overstates weak presence, so pair it with a per-answer rubric that rewards the quality of inclusion:
- 3 points: recommended first or named as the best fit for the stated need.
- 2 points: named with accurate, positive framing.
- 1 point: listed neutrally among alternatives.
- 0 points if absent; −1 point if named inaccurately or negatively. Add +1 to any non-negative score when an owned URL is cited.
Normalize to 0–100 by dividing total points by the maximum possible (4 × answers) and multiplying by 100. The score gives you one trend line, and the component metrics explain why it moved.
LLM visibility tutorial: how to measure it step by step
You can measure LLM visibility with a prompt library, a fixed sampling schedule and a spreadsheet before buying any tool. Commercial platforms automate the same method, so running it by hand once also tells you what to demand from a vendor. Our piece on AEO tracking covers the version focused on answer engines.
- Build a library of 30–100 prompts. Split it by funnel stage (problem-aware, category, comparison, vendor evaluation), intent and product line. Keep branded prompts ("Is [brand] good for SOC 2 evidence collection?") apart from non-branded ones ("best CSPM tools for a 500-person company"), and weight non-branded prompts heavily, since that is where shortlists form.
- Define your entity list. Record brand and product names, former names, common misspellings and acquired brands for you and 4–8 competitors. Without alias matching, "Acme Cloud" and "AcmeSec" count as two brands, and a generic phrase such as "endpoint protection" can be mistaken for a product.
- Fix engines and markets. Pick the engines your buyers use, typically ChatGPT, Perplexity, Gemini and Google AI Overviews, and set the region and language for each run.
- Run each prompt 3–5 times per engine on a fixed weekly schedule, in clean sessions with no logged-in history.
- Capture every answer in full, not just a yes/no for mentions.
- Parse and score. Tag mentions against the entity list, record mention order, classify sentiment and label each cited domain as owned, competitor, review site, Reddit, media or reference.
- Normalize and report per engine, per prompt group and overall. Weight groups by business value so a cluster of 40 informational prompts does not outweigh 10 comparison prompts.
Minimum data record per answer
- Prompt ID, prompt text, funnel stage and branded/non-branded flag.
- Engine, model version where exposed, collection method (API or browser) and whether web search was active.
- Timestamp, region and language.
- Full answer text, brands mentioned in order and sentiment label.
- Every cited URL and its domain category.
A worked example
Take a mid-market cloud security vendor running 40 prompts across ChatGPT, Perplexity, Gemini and Google AI Overviews, five runs each, for 800 answers a week. Illustrative results show a 31% overall mention rate, split as 44% on Perplexity and 22% on ChatGPT. On the 12 non-branded comparison prompts, the mention rate falls to 12%, and two competitors appear in more than 60% of those answers. Owned pages make up 4% of citations. Competitor comparison pages, G2 listings and two Reddit threads make up most of the rest.
The citations explain the gap. The vendor has no comparison pages of its own, so engines quote the competitors' versions. It is thin on review sites, and it is absent from the Reddit threads Perplexity and AI Overviews keep citing. So the vendor should publish three comparison pages and answer-shaped articles for the comparison prompts, run a review-generation push and join the cited Reddit threads with disclosed, useful replies. Re-measure the same 40 prompts after six weeks so the before-and-after comparison uses an identical sample.
Why one-off checks and dashboard scores mislead
AI answers are non-deterministic, so a single screenshot is one draw from a distribution rather than your actual visibility. Models sample tokens with some randomness, engines decide per query whether to search the web, and providers update models without notice. Security teams testing LLM systems reached the same conclusion: organizations deploying LLM-powered systems need continuous visibility, not a point-in-time audit.
Sample size matters too. With 200 answers and a 30% mention rate, the 95% margin of error is roughly ±6.4 percentage points (1.96 × √(0.3 × 0.7 ÷ 200)), so a week-over-week move from 30% to 26% is noise. Report rates with their sample size, and treat a move as real only if it exceeds the margin or persists for two to three weeks.
What else distorts the numbers
- Model drift: a provider update can shift answers across your whole prompt set overnight with no change on your side.
- Localization: the same prompt returns different brands in the US, UK and Germany, and in different languages.
- Uncited answers and intermediaries: answers from training data leave only the mention to track, and engines often cite a listicle, review site or Reddit thread that summarizes your page, so citation rate understates your influence.
- Hallucination: a model can invent features, pricing or integrations, and a mention counter scores that as a win unless someone checks accuracy.
Engines retrieve information differently enough that blended averages hide the signal:
| Surface | Retrieval behavior | Measurement implication |
|---|---|---|
| ChatGPT | Answers from training data or triggers web search, depending on the query | Track whether search was used; score cited and uncited answers separately |
| Perplexity | Searches the web on nearly every query and cites sources inline | Most citation-rich surface; best for diagnosing which sources drive inclusion |
| Gemini | Can ground answers in Google Search | Visibility tends to follow Google index strength |
| Claude | Answers from training data unless web search is enabled | Mostly a test of what the model already knows about the brand and how clearly it identifies the entity |
| Copilot | Grounds answers in Bing results | Bing indexing and rankings matter |
| Google AI Overviews and AI Mode | Built on Google's index, with linked sources beside the summary | Closest to SEO; track cited domains per query alongside organic results |
Report every metric per engine first, then blend. A 10-point gain on Perplexity and a 10-point loss on ChatGPT average to "flat" and hide two separate stories.
How to read the numbers and turn them into action
Start by separating content problems, where models do not know or trust you, from retrieval problems, where they cannot reach your pages or prefer other pages about you. The pattern of a drop tells you which one you have.
- Drop on one engine, across many prompts: usually a model or retrieval change. Check crawler access first, since a robots.txt rule or CDN bot setting blocking GPTBot, OAI-SearchBot or PerplexityBot cuts you out of those engines' live retrieval.
- Mentions hold, citation share falls: competitors or intermediaries published fresher, more quotable pages. That is a content gap on specific prompts.
- Drop across every engine, including uncited answers: a reputation or entity problem, such as negative reviews, a confusing rebrand or a competitor dominating third-party coverage.
- Rising mention rate, falling sentiment: you are discussed more, but in complaint threads or unfavorable comparisons.
A dashboard reports that answer share fell 8 points. A diagnosis names the prompts, engines and cited pages that replaced you, and the action list follows from those pages. As of October 2026, Profound's G2 reviewers ask for exactly this, wanting more context on why scores change rather than only that they did. Measurement produces action only when each cadence has an owner:
| Cadence | Work | Owner |
|---|---|---|
| Weekly | Run the prompt set, check per-engine rates, flag moves beyond the margin of error and inaccurate answers | SEO |
| Monthly | Review cited sources and displacement; assign new pages, refreshes and review or PR outreach | Content and PR |
| Quarterly | Refresh the prompt library, update competitor and alias lists, re-weight prompt groups to pipeline priorities | Product marketing |
Manual, in-house or commercial: choosing how to measure
The choice depends on how many prompts you track, how many engines and markets you cover, and whether you need the work to produce fixes or only to report.
- Manual: suits a first baseline of 20–40 prompts on two engines. Beyond that, run counts outgrow what one person can capture consistently.
- In-house pipeline: calling model APIs gives teams with data engineers full control over sampling and scoring, much as Google's engineers built an evaluation framework to test LLM output systematically. API answers can differ from what logged-in users see, and someone must maintain parsing and entity resolution.
- Commercial platforms: handle collection at scale. Our roundup of LLM visibility tools for ChatGPT and Claude covers the wider field.
Validate any platform before trusting it. Request raw prompts, full answer text, timestamps, cited URLs, region and language metadata, model version and collection method for a sample, then rerun ten prompts yourself and compare.
| Option | Surfaces and collection | Metrics and refresh | Produces the fix | Pricing model |
|---|---|---|---|---|
| Tellr | Google organic results, AI Overviews and discussions; own search-data pipeline | Weekly citation share by domain category, Reddit placements, content rankings and citations | Yes: content, Reddit replies, ad creative | Managed engagement, scoped on a call |
| Semrush | ChatGPT, Perplexity, Gemini, Claude, Copilot, AI Overviews, AI Mode; mainly API | Mentions, citations, share of voice, position; prompt rankings daily, brand data weekly | No: insights only | Per domain |
| Profound | Same seven surfaces; API, browser sessions and panel data | Adds sentiment and AI referral traffic; index weekly | Yes: briefs, drafts, page optimization | Custom enterprise |
| Scrunch AI | Same seven surfaces; API, browser sessions and panel data | Mentions, citations, share of voice, sentiment, position, AI bot traffic | No: reporting and tracking | Tiered by prompt volume |
Tellr
Tellr is a managed earned-visibility program rather than a self-serve tracker. Each week it tracks the category's queries on Google and records the organic results, the AI Overview and the discussions block. It labels each cited domain as your site, a competitor, Reddit, social, review sites or references, and shows week-over-week movement.
- Writes comparison pages, reviews and answer-shaped articles for ChatGPT, Perplexity, Gemini, Claude and AI Overviews to quote, and publishes them to WordPress.
- Places Reddit replies in the threads engines cite and reports whether each is still live.
- Applies a brand brief, claim guardrails, an approval gate and an audit trail to everything placed.
- Offers API access and an MCP server, "Tellr for Claude", for plain-language questions about the program.
Semrush
Semrush's AI Visibility Toolkit suits SEO teams that want AI tracking beside their existing keyword and rank data. It costs about $99 a month per domain billed annually, plus $45 a month per extra user and $60 a month per 50-prompt pack. According to G2, Semrush holds a 4.5 rating from 3,945 reviews as of October 2026.
- Daily prompt-ranking data across seven AI surfaces.
- Integrates with GA4, Search Console, WordPress, an API and an MCP server.
- Reviewers cite limited regional and language coverage and caps on prompts and exports.
- Sentiment is inconsistently documented, and the tool reports problems without producing fixes.
Profound
Profound suits enterprise teams that want deep multi-engine data and content workflows in one platform, with SSO, approval workflows, an audit trail and SOC 2 Type II. Pricing is custom. A 7-day trial includes 50 prompts a day across ChatGPT, Gemini and AI Overviews. According to G2, it holds a 4.6 rating from 1,124 reviews as of October 2026.
- Mixed API, browser and panel collection.
- Generates briefs, drafts and page optimizations from the data.
- G2 reviewers report a steep learning curve, limited export flexibility and "Average Position Rank" results that surface generic non-brand entries. That is an entity-resolution issue worth testing.
- Some users question the reliability of scraping-based collection.
Scrunch AI
Scrunch AI, acquired by Sitecore in June 2026, suits agencies and mid-size teams that want custom prompt tracking at published prices. Starter costs $300, Growth $500 and Pro $1,000 a month, with a 7-day trial. According to G2, it holds a 4.6 rating from 73 reviews as of October 2026.
- Citation tracking, competitor comparison and GA4 integration.
- AI bot and agent traffic monitoring.
- Reviewers report inconsistent mention counts between views, limited historical trend data and unclear refresh timing.
- Prompt credits run down fast because each engine counts separately, and the tool stops at reporting.
Where Tellr fits in an LLM visibility program
Tellr fits teams that already know where they are missing and need a governed program to close the gap, not another report. The workflow runs in four stages:
- Weekly citation data shows which domains Google and AI Overviews quote for your category's queries.
- The senior team turns those gaps into briefs, drawing on your knowledge base and real user reviews.
- Pages, Reddit replies and creative pass claim checks and your approval gate before going live.
- A weekly digest and a monthly program review show what shipped and whether it ranks and gets cited.
Tellr is built for marketing teams spending $10k+ a month at companies worth $500M+ or with 200+ employees. A smaller team that only needs tracking will get better value from Semrush or Scrunch AI's lower tiers.
A measurement checklist to start this quarter
Over one quarter, you can set up a fixed sample, honest statistics and a named owner for every action:
- Week 1: write the prompt library and the entity and alias list.
- Week 2: run the first full sample and record a per-engine baseline with sample sizes and rubric scores.
- Weeks 3–6: diagnose by engine pattern, check crawler access, and ship the pages, reviews and replies the citations point to.
- Week 6 onward: re-measure the identical prompt set and confirm owners for each weekly, monthly and quarterly cadence.
By quarter end you will have a baseline, a trend line and a list of the pages and sources standing between you and the shortlist. Use your LLM visibility measurement to decide what to publish, fix and earn next.
FAQ
What is LLM visibility?
LLM visibility measures how often, how prominently, and how accurately AI systems such as ChatGPT, Gemini, Claude, Perplexity, Copilot, and Google AI Overviews mention, recommend, or cite your brand in generated answers.
Is LLM visibility the same as a ranking?
No. LLM visibility is not a fixed ranking because AI answers are generated fresh each time. The same prompt can include your brand in one run and exclude it in another, so visibility is measured as rates across repeated samples rather than positions alone.
How do you measure LLM visibility accurately?
Use a fixed prompt library split by funnel stage, branded and non-branded intent, and engine. Run each prompt 3 to 5 times per engine on a fixed schedule, capture every full answer, then calculate metrics such as mention rate, citation share, answer share, sentiment, and prompt coverage.
Why are one-off checks unreliable?
A single check is only one draw from a non-deterministic system. AI answers vary because of randomness, model updates, localization, and whether the engine uses web search, so one screenshot cannot represent your true visibility.
What does a drop in LLM visibility usually mean?
A drop on one engine usually points to a retrieval or model change, while a drop across every engine usually signals a content, reputation, or entity problem. If mentions hold but citation share falls, competitors or intermediaries may have published fresher or more quotable pages.