A ranking tells you where a page sits in a list of links. An AI visibility score tells you whether the answer a buyer reads names you at all, and how it describes you. The score is a composite metric, usually on a 0–100 scale, that measures how often and how prominently a brand appears in answers from AI engines such as ChatGPT, Perplexity, Gemini and Google AI Overviews, across a fixed set of buyer prompts. It counts mentions, citations and position, then weights them by prompt intent and engine. Each score is a dated reading. The engines are asked real questions, the answers are archived, and the score counts what came back. It does not predict anything. This guide, current as of October 2026, covers how the score is calculated, where the method breaks, what good looks like against competitors and over time, and which fixes move it fastest.
Key takeaways
- An AI visibility score is a sampled, weighted measure of brand presence in AI answers, not a ranking and not a business outcome.
- The prompt set decides an AI visibility score more than any other input, so a score is only comparable when the prompts, engines and locations stay fixed.
- A good AI visibility score beats the competitors buyers compare you with, rises over time, and comes from non-branded category prompts.
- Being mentioned, being cited and being recommended are separate outcomes, and only recommendation in high-intent prompts reliably reflects buyer influence.
- Trend lines across several reruns are more reliable than any single snapshot, because AI answers vary by day, device, location and model version.
Why AI visibility needs a different measurement model than SEO
AI engines return one synthesized answer instead of a ranked list of links, so position-based metrics cannot tell you whether your brand is named. Search engine optimization is the practice of improving the visibility and performance of web pages in search engine results. Its metrics assume ten blue links, a click and a landing page. An AI answer breaks all three assumptions:
- There is no fixed position. A brand is named first, named fifth or absent.
- The cited source is often not the brand. A Reddit thread or review site can earn the citation while the brand earns the mention.
- The same prompt returns different answers on different days, devices and accounts.
- The buyer may never click. The answer itself shapes the shortlist.
That is why LLM visibility belongs next to SEO reporting rather than inside it. The table shows where each metric fits.
| Metric | What it measures | What it misses | Best use |
|---|---|---|---|
| AI visibility score | Weighted presence of the brand in AI answers across a prompt set | Clicks, revenue, the exact wording each buyer sees | Tracking share of AI answers against competitors over time |
| Organic rankings | Position of a URL for a keyword | Whether an AI Overview above the results names you | Page-level SEO performance |
| Local pack visibility | Presence in map results for local queries | Conversational and comparison prompts | Location-based businesses |
| Share of voice | Your mentions as a share of all brand mentions | Sentiment and whether you are recommended | Competitive benchmarking |
| Citation share | Your domain's share of all sources cited | Mentions that carry no link | Judging whether your own pages are trusted sources |
| AI referral traffic | Sessions arriving from AI engines | Influence on buyers who never click | Downstream evidence, read alongside the score |
What an AI visibility score measures, and what it does not
The score counts how often, how prominently and how favourably a brand appears in sampled AI answers. It says nothing about traffic, pipeline or what any individual buyer saw. It is a sample of answers, and every number in it depends on which prompts were asked, where, on which engine and on which day.
Mentioned, cited and recommended
Most scores collapse three different outcomes into one number. Separate them before you read anything else.
- Mentioned: the brand name appears in the answer, perhaps in passing or in a list of ten.
- Cited: the engine links to a source. That source may be your domain, a competitor's comparison page or a Reddit thread that discusses you.
- Recommended: the answer presents the brand as a good choice for the stated need, often first or with a reason attached.
A security vendor named as "another option" in a list of eight gets the same mention count as one recommended first "for multi-cloud teams". Only the second influences a shortlist. Teams serious about tracking whether AI answers mention you log all three outcomes separately.
The dimensions behind the number
| Dimension | Input logged per answer | Failure state | Typical remediation |
|---|---|---|---|
| Presence | Brand named: yes or no | Absent from non-branded category prompts | Answer-shaped pages and third-party coverage for those prompts |
| Prominence | Order in the answer; recommended or listed | Named late, or only as an alternative | Comparison pages that state clear use-case fit |
| Citation | Domains cited and which passage was used | Mentioned but never cited; competitors own the cited pages | Pages with quotable, specific passages; crawler access |
| Accuracy and sentiment | Positive, neutral, negative or factually wrong | Outdated pricing, wrong features, negative framing | Correct the source pages and the threads engines draw from |
The score does not tell you revenue impact, how many buyers asked those prompts, or why an engine chose a source. It is a diagnostic, and acting on the diagnosis is a separate job.
How an AI visibility score is calculated
Run a fixed set of buyer prompts across several AI engines, record whether each answer mentions, cites or recommends the brand, and convert the weighted results into a percentage or 0–100 figure. The simplest version is:
Visibility score = (AI responses mentioning the brand ÷ total AI responses analyzed) × 100
Most methods then add share of voice (your mentions over all brand mentions), average position in list-style answers, citation share (your domain's citations over all citations) and a framing weight for answers that present the brand as an authority.
A weighted formula you can reproduce
The model below is illustrative. Score each answer with a points rubric:
- Recommended first or with an explicit reason: 3 points
- Mentioned without recommendation: 1 point
- Own domain cited: +1 point
- Factually wrong or negative framing: 0 points, flagged for review
Then: Score = Σ (prompt weight × engine weight × answer points) ÷ maximum possible points × 100. For example, weight high-intent comparison prompts at 2 and informational prompts at 1. If 60 prompts run across 4 engines, the score reflects 240 answers, and a single answer swing moves the result only slightly. A composite score also rarely equals the simple average of its pillars, because the weights pull it toward the prompts and engines you rated as most important.
Prompt design
The prompt set decides the score. Branded queries flatter you, and vague informational questions measure nothing commercial. Cover each of the eight intents your buyers have.
| Prompt type | What it tests | Example (cloud security buyer) |
|---|---|---|
| Discovery | Whether you make the category shortlist | "What are the best CSPM tools for AWS and Azure?" |
| Comparison | How you are framed against named rivals | "Wiz vs Orca vs Prisma Cloud for a 2,000-person company" |
| Urgency | Presence when a problem is live | "How do I find exposed S3 buckets right now?" |
| Trust | Reputation and risk signals | "Is [vendor] reliable for regulated industries?" |
| Price | Value framing and pricing accuracy | "Most cost-effective cloud security platform for mid-market" |
| Brand substitution | Whether you appear as an alternative to a leader | "Alternatives to [market leader]" |
| Informational | Whether your content is the cited explanation | "What is CNAPP?" |
| Local intent | Regional or location-specific presence | "Cloud security consultancies in Frankfurt" |
Quality controls for the prompt set:
- Version the set and freeze it between reruns; add new prompts as a separate cohort.
- Write two or three phrasings per intent so one wording does not decide a cluster.
- Split branded and non-branded prompts and report them separately.
- Take prompts from real buyer language: sales call notes, Reddit threads, Search Console queries.
Engines, collection method and rerun cadence
Collection method changes results, so know which one sits behind your number.
- API calls return model output without the consumer app's personalization or, in some cases, its live web retrieval.
- Browser sessions capture what a consumer sees but vary with location and logged-in state.
- Panel data reflects real users but in smaller, less controlled samples.
To normalize across engines, score each engine separately first, then combine the results using weights you choose and document, such as your estimate of where your buyers ask. Never average raw mention counts across engines that answered different numbers of prompts.
Rerun cadence should follow volatility. Fast-moving categories such as AI software or consumer security, or any category during a launch, justify weekly runs. Stable B2B categories can run every two to four weeks. Run each prompt more than once per cycle and record the spread of answers, since a single answer can mislead.
The AI Visibility Score in Semrush and other vendors
The AI Visibility Score in Semrush is a 0–100 benchmark of how often and how clearly a brand appears in AI-generated answers across ChatGPT, Gemini, AI Mode and AI Overviews, read relative to competitors. According to G2, Semrush holds 4.5 out of 5 from 3,945 reviews as of October 2026, and reviewers increasingly single out its AI Overview and AEO tracking. Vendors differ mainly in which surfaces they sample, how they collect answers and what they report.
| Vendor | Surfaces | Collection | What it reports | Refresh |
|---|---|---|---|---|
| Tellr | Google organic results, AI Overviews, discussions block | Own search-data pipeline | Citation share by domain type (own site, competitor, Reddit, social, review sites, references), week-over-week movement | Weekly, plus monthly program review |
| Semrush | ChatGPT, Perplexity, Gemini, Claude, Copilot, AI Overviews, AI Mode | Mainly API, some browser-based | 0–100 score, mentions, citations, share of voice, position | Daily prompt rankings, weekly brand and competitor data |
| Profound | Same seven surfaces | API, browser sessions and panel data | Mentions, citations, share of voice, sentiment, position, AI referral traffic | Core index weekly |
| Peec AI | ChatGPT, AI Overviews, AI Mode, Perplexity, Gemini, Copilot; Claude as add-on | Mainly browser emulation, API for some models | Mentions, URL-level citations, share of voice, sentiment, position; no referral traffic | Daily |
Reviewers raise different trade-offs for each tracker.
- Semrush: reviewers complain about limited regional and language coverage, and say it is more monitoring than workflow; its per-domain pricing suits smaller teams well.
- Peec AI: praised for simple tracking, criticized for stopping at monitoring.
- Profound: seen as deep but costly, with data that can be hard to turn into next steps.
How to audit your score manually
You can approximate a score with a spreadsheet before buying anything. ZDNet's guide covers how to check whether ChatGPT and other AI tools cite your website, and the process below extends it into a repeatable score.
- Write 30–50 prompts across the eight intent types, split branded and non-branded.
- Run each prompt in a logged-out or clean session on each engine, with a fixed location.
- Archive every answer as text and screenshot, with date, engine and model version.
- Log per answer: mentioned, position, recommended, sentiment, and each cited URL.
- Extract cited domains and label them as own, competitor, Reddit, review site or reference.
- Apply the points rubric and compute the score per engine, then the weighted composite.
- Rerun the same set two or three times in the first week to measure natural variance.
Editor's tip: The variance you measure in step 7 is your noise floor. If three runs of the same set land at, for example, 38, 41 and 44, a later move from 41 to 43 is not a result worth reporting.
For a structured version, see how to run an AI visibility audit in a week.
What a good AI visibility score looks like
No single number is good in every category. A 35 in a crowded category with ten credible vendors can be stronger than a 70 in a niche with two. A good AI visibility score beats the competitors buyers compare you with, rises across successive reruns, and comes from non-branded category prompts rather than your own brand name. Read the score through relative benchmarks:
- Competitor quartile: score your top five to ten competitors on the same prompt set and see which quartile you sit in.
- Trend: compare against your own baseline; treat movement inside your measured rerun variance as noise.
- Branded vs non-branded: branded presence should be near-total; the real target is non-branded discovery and comparison prompts.
- Prompt-win rate: the share of high-intent prompts where you are recommended first or second.
- Citation ownership: how often the cited passage comes from your domain or a source that represents you accurately.
A maturity model
These five bands are a working framework for illustration and are not drawn from industry data. Calibrate them to your category.
- Invisible: absent from most non-branded prompts; competitors and directories own the answers.
- Emerging: mentioned in some discovery prompts, rarely recommended, seldom cited.
- Competitive: named in most category prompts, recommended in some comparison prompts, mid-pack against rivals.
- Leading: top-quartile against competitors, frequently recommended first, own pages cited regularly.
- Dominant: recommended across nearly every prompt cluster and engine, with your passages quoted as the reference explanation.
Worked examples: same score, different meaning
All three scenarios use illustrative figures, and each brand scores 42.
| Scenario | What sits behind the 42 | Priority |
|---|---|---|
| B2B cloud security vendor | Competitors score 30–55, so it is mid-pack. 95% presence on branded prompts, 31% on non-branded discovery, recommended first in 4 of 20 comparison prompts. Perplexity cites its comparison pages; Google AI Overviews cite Reddit threads where it is barely discussed. | Close the Reddit gap before writing more blog posts. |
| Consumer security app | The leader scores 78. Informational prompts ("what is a VPN") inflate the score because its glossary is cited, while it is missing from "best VPN for streaming" price and comparison prompts. | Comparison and price content; this 42 is weaker than the B2B case. |
| Multi-location clinic group | The national score hides a split: 70 in two metro areas with strong review profiles, under 15 in six others. Local-intent prompts run from each city's location settings expose it. | Location pages and review coverage in each market, ahead of national content. |
What causes false confidence
- High branded visibility masking low category visibility.
- Strong results on one engine only, often the one whose answers rely most on your own site.
- Presence carried by directories or aggregator listings you do not control.
- Being cited without being recommended, such as a glossary page quoted for a definition.
- Positive mentions that contain outdated pricing or features.
Industry matters too. Ecommerce answers lean on reviews and price data. SaaS answers lean on comparison pages and community threads. Healthcare and legal answers favour authoritative and regulated sources, and engines hedge more. Enterprise brands with local branches need national and local prompt sets scored separately.
How to improve an AI visibility score
Change the sources AI engines draw from: your own answer-shaped pages, the third-party threads and reviews that discuss you, and the technical access engines need to read your site. Work through them by signal type.
Owned content
Engines quote passages that answer a question directly and specifically. Comparison pages, use-case pages and articles that open each section with a self-contained answer get quoted more than broad thought-leadership posts. State facts engines can lift: who the product suits, what it integrates with, and how it differs from named rivals.
Third-party corroboration
Google AI Overviews and Perplexity often cite Reddit, review sites and industry publications. If buyers discuss your category in r/sysadmin or r/cybersecurity and you are absent, the answer reflects that. Genuine, guideline-compliant participation and review coverage change the source pool. Bought upvotes and fake reviews create risk, and moderators and platforms routinely remove them.
Technical access
Check robots.txt and your CDN bot rules for the crawlers each engine uses:
- OAI-SearchBot: retrieval for ChatGPT search.
- PerplexityBot: Perplexity's index.
- ClaudeBot: Anthropic's crawler.
- Googlebot: used for AI Overviews; blocking Google-Extended affects Gemini training use, not Overviews.
Serve key content in server-rendered HTML rather than client-side JavaScript, add accurate schema such as Organization, Product and FAQPage, and keep pricing and feature pages current.
| Issue | Symptom in score | Fix | Effort | Typical time to move |
|---|---|---|---|---|
| AI crawlers blocked | Never cited on one engine | Update robots.txt and CDN bot rules | Low | Weeks, once recrawled |
| Outdated facts in answers | Negative or wrong framing | Correct source pages; publish dated updates | Low | Weeks to a month |
| No comparison content | Absent from comparison and substitution prompts | Publish honest comparison and alternatives pages | Medium | One to three months |
| Absent from community threads | Competitors cited via Reddit in AI Overviews | Sustained, compliant Reddit participation | Medium to high | A quarter or more |
| Weak review footprint | Low trust-prompt and local-prompt presence | Review programs on the sites engines cite | Medium | A quarter or more |
How Tellr runs AI visibility as a program
Tellr uses measurement to steer the work and does not sell the measurement as a product of its own. Its weekly citation tracking shows which domains Google and its AI Overviews cite for the category's queries, and a senior team then acts on those gaps in one governed program, with guardrails, approvals and an audit trail. Tellr is a managed program for mid-market and enterprise marketing teams. A smaller team that wants a dashboard to run itself will be better served by a self-serve tracker.
- Citation gaps turned into comparison pages, reviews and answer-shaped articles written to be quoted by ChatGPT, Perplexity and AI Overviews, published to your CMS.
- Subreddit mapping, a daily thread radar and guideline-checked Reddit replies behind an approval gate.
- Paid-media ad intelligence and ready-to-run creative in the same program.
- Weekly reporting and a monthly program review that tie score movement to the work shipped.
Limitations, bias and quality controls
The score is a sample shaped by the choices behind it, so its main risks are biased prompt sets, uncontrolled collection settings and inconsistent extraction rules. Know where each bias enters:
- Prompt selection and phrasing: small wording changes can flip which brands appear.
- Location, personalization and logged-in state: the same prompt answers differently by city, account history and session.
- Freshness and model updates: retrieval sources and model versions change, sometimes overnight.
- Citation extraction and deduplication: tools differ on whether a subdomain, a product name variant or a repeated mention counts once or several times.
- Hallucinated mentions and citation mismatch: engines can name you with invented features, or cite a page that does not support the claim.
- Sentiment: a disqualifying mention ("not suitable for enterprises") should not add to your score.
These quality controls keep the number honest:
- Freeze and version prompts, engines, locations and the points rubric.
- Run each prompt several times per cycle and report the range alongside the score.
- Annotate the trend line with model releases and your own launches.
- Spot-check a sample of logged answers by hand each cycle.
- Read score movement next to AI referral traffic, branded search demand and sales-reported "found you via ChatGPT" mentions, without claiming direct attribution.
Is the score worth reporting to leadership? Yes, as a trend with its range shown and annotated with what changed. A single reading without the rerun spread invites decisions based on noise.
Treat the first run as a baseline to measure against. An AI visibility score earns its place in reporting when the method stays fixed, competitors are scored the same way, and the trend is read against what changed in your content, your community presence and the engines themselves.
FAQ
What is an AI visibility score?
An AI visibility score is a composite 0–100-style metric that measures how often and how prominently a brand appears in AI answers across a fixed set of buyer prompts. It tracks outcomes such as mentions, citations and recommendations, then weights them by prompt intent and engine.
How is an AI visibility score different from a ranking?
A ranking shows where a page sits in a list of links. An AI visibility score shows whether the AI answer names your brand at all, whether it cites your domain, and how it frames you. It is designed for synthesized answers, not traditional search result positions.
What does a good AI visibility score look like?
A good score beats the competitors buyers actually compare you with, improves across repeated reruns, and is driven by non-branded category prompts rather than your own brand name. The article stresses that no single number is universally good across every category.
Why are trend lines more useful than a single AI visibility score?
AI answers vary by day, location, device, account state and model version, so one reading can be noisy. Repeated runs on the same prompt set reveal the normal variance and make it easier to separate real movement from ordinary fluctuation.
How can you improve an AI visibility score?
The fastest gains usually come from improving the sources AI engines draw from: publish answer-shaped comparison and use-case pages, strengthen third-party coverage such as reviews and community discussion, and make sure AI crawlers can access key content in server-rendered HTML with accurate schema and current facts.