AI visibility is how often AI-generated answers in ChatGPT, Perplexity, Gemini, Claude, Copilot and Google AI Overviews mention, describe, cite or recommend your brand when buyers ask about your category. For most enterprise brands, the measured result is uneven. They are named reliably on branded prompts and rarely on the non-branded, high-intent prompts where shortlists form. This guide sets out a repeatable audit method for October 2026. It covers how to define the metrics, build and rerun a prompt set, score answers against three to five competitors, diagnose why the brand is missing, and turn the results into monthly reporting.
Key takeaways
- AI visibility is five separate metrics (mention rate, recommendation rate, citation frequency, sentiment and positioning, and share of AI voice), and a brand can lead on one while failing another.
- A credible AI visibility audit reruns every prompt three to five times per engine under fixed conditions, because single answers vary too much to report.
- Non-branded category and high-intent buyer prompts matter more than branded prompts, because that is where AI answers build shortlists.
- Brands that rank well on Google can still be missing from AI answers when third-party sources such as review sites, Reddit threads and trade press do not name them.
- Depth of description and presence in comparison prompts usually improve before recommendation-level mentions appear, so they are the leading indicators to track monthly.
What AI Visibility Actually Measures
AI visibility measures whether an AI answer includes your brand, how it describes the brand, and whether it recommends the brand over competitors. Treating it as one number hides the problem. Being named in 60% of answers means little if the brand is recommended in 5% of them.
| Metric | What it measures | What it tells you |
|---|---|---|
| Mention rate | Share of answers that name the brand | Whether the model recognizes you as part of the category |
| Recommendation rate | Share of answers that suggest the brand as a fit or top pick | Whether you make the shortlist |
| Citation frequency | Share of cited answers that link to a brand-owned URL | Whether your pages are used as evidence |
| Sentiment and positioning | Tone, accuracy and the use case the answer assigns you | Whether the model describes you the way you sell |
| Share of AI voice | Your mentions as a share of all tracked brands' mentions | Your competitive standing in the category |
AI visibility depends on work most teams already do:
- SEO: strong rankings help AI Overviews find your pages, but rankings alone do not guarantee inclusion.
- Brand search: branded queries often rise once assistants start naming you.
- Review sites: answers about software often cite G2, Capterra and Gartner Peer Insights pages.
- PR and thought leadership: trade press, expert quotes and original data give models independent sources that agree about you.
This matters more now because buyers form shortlists inside answers instead of on a results page. Forrester's guide to answer engine optimization warns that content not built for AI search leaves a brand invisible to the buyers who use it.
Patterns also differ by category. In cybersecurity, AI is now part of the product itself. 77% of security stacks include AI, according to the Cloud Security Alliance's State of AI Cybersecurity 2026, which also finds that trust is lagging. That makes credible third-party evidence about each vendor more important.
| Industry | Typical AI mention pattern | Sources answers tend to lean on |
|---|---|---|
| B2B SaaS | Long lists of named options, with recommendations split by company size | Review platforms, comparison pages, Reddit threads |
| Professional services | Few firm names; answers often describe criteria rather than providers | Directories such as Clutch, rankings, trade press |
| Healthcare | Models hedge and often decline to recommend a single provider | Institutional sources such as NIH and major clinics |
| Cybersecurity | Incumbents dominate category prompts; challengers appear in comparisons | Analyst coverage, practitioner subreddits, peer reviews |
| Ecommerce | Product-level answers with prices and availability | Merchant Center feeds, product reviews, Reddit |
Takeaway: decide which of the five metrics you are reporting before you measure anything.
How to Run a Repeatable AI Visibility Audit
A repeatable AI visibility audit uses a fixed prompt set, controlled run conditions and several reruns per prompt, so that changes between months reflect real movement rather than noise.
- Build 40 to 60 prompts across four types. Use branded prompts ("What does [Brand] do?"), non-branded category prompts ("best cloud security platforms"), comparison prompts ("[Brand] vs [Competitor]") and high-intent buyer prompts ("CNAPP for a regulated bank running AWS and Azure").
- Vary the prompts by funnel stage, buyer role and geography. A CISO asking about compliance coverage gets a different answer from a DevOps lead asking about agent overhead. A prompt run from Germany can surface different vendors than the same prompt run from the US.
- Fix the run conditions. Use the API or logged-out sessions, turn memory and custom instructions off, set the locale with a VPN, and record whether browsing or web search was on.
- Rerun each prompt three to five times per engine within the same 48-hour window, and log the model version and the date.
- Log every answer verbatim. Record the prompt ID, engine, run number, mode, brands named in order, cited URLs and any factual errors.
- Normalize scores per engine. Calculate rates within each engine first, then average them. Pooled raw counts overweight whichever engine you ran most.
| Dimension | Example prompt |
|---|---|
| Awareness stage | "How do companies secure workloads across multiple clouds?" |
| Evaluation stage | "Which CNAPP vendors are best for 5,000+ workloads?" |
| Buyer role: CFO | "How is cloud security software usually priced for enterprises?" |
| Geography: EU | "GDPR-ready cloud security platforms for a German retailer" |
Limitations to document in every report: models hallucinate features and pricing, the same prompt gives different answers on reruns, results shift by region, browsing and non-browsing modes draw on different sources, and a model update can change answers overnight. Flag any month that spans a major model release instead of comparing it directly with the month before.
Takeaway: a smaller prompt set that you rerun consistently beats a large set that you run once.
How to Score Answers and Benchmark Competitors
Score every logged answer on a fixed rubric, then calculate the same metrics for your brand and three to five direct competitors from the same answers. The scoring rubric has two scales:
- Depth (0 to 3): 0 means not mentioned, 1 means name only, 2 means one accurate descriptor, and 3 means accurate differentiators and use case.
- Recommendation strength (0 to 3): 0 means absent or negative, 1 means listed neutrally, 2 means recommended among options, and 3 means the primary pick.
Use these formulas for the core metrics:
- Mention rate = answers naming the brand ÷ total answers
- Recommendation rate = answers with a recommendation score of 2 or more ÷ total answers
- Citation frequency = answers citing a brand-owned URL ÷ answers with any citation
- Average placement = mean position among the brands listed, counted only in answers that name the brand
- Share of AI voice = brand mentions ÷ total mentions of all tracked brands
To measure share of voice in other channels as well, see our guide to measuring share of voice across search, social and AI.
Worked example: a cloud security audit
The following example uses illustrative figures. Brand A, a cloud security vendor, ran 40 prompts × 3 runs × 4 engines (ChatGPT, Perplexity, Gemini and Google AI Overviews), for 480 answers. Here are two raw outputs and how each was scored:
- Perplexity, "best CNAPP for a multi-cloud enterprise": "Leading options include Competitor B, known for agentless scanning, and Competitor C… Brand A is another option." This scores placement 3, depth 1 and recommendation 1. The answer cites G2, an r/cloudsecurity thread and Competitor B's comparison page.
- ChatGPT, "Brand A vs Competitor B for AWS and Azure": the answer describes Brand A's runtime protection accurately and recommends it for teams that prioritize runtime detection. This scores depth 3 and recommendation 2.
| Brand | Mentions (of 480) | Mention rate | Recommendation rate | Avg placement | Share of AI voice |
|---|---|---|---|---|---|
| Competitor B | 290 | 60% | 31% | 1.8 | 33% |
| Competitor C | 220 | 46% | 18% | 2.6 | 25% |
| Brand A | 168 | 35% | 10% | 3.4 | 19% |
| Competitor D | 130 | 27% | 9% | 3.9 | 15% |
| Competitor E | 72 | 15% | 4% | 4.5 | 8% |
Interpretation: Brand A's mention rate by prompt type is 92% on branded prompts, 58% on comparison prompts, 21% on category prompts and 9% on high-intent prompts. In 2 of 12 runs, Gemini wrongly stated that Brand A lacks Azure support. The brand is recommended mainly when a buyer already names it.
Prioritization:
- Correct the Azure error on brand-owned pages and on review profiles.
- Get Brand A into the Reddit threads and G2 comparisons that the category answers cite.
- Publish answer-shaped "best CNAPP for [use case]" and comparison pages.
Takeaway: the gap between comparison prompts and category prompts usually points to the fix.
Why AI Answers Mention Some Brands and Skip Others
AI answers mention brands that models can identify as distinct entities and that several independent sources describe in the same way. Content structure matters, but it is only one of these factors:
- Entity recognition: use one consistent brand name, Organization schema with sameAs links, and Wikidata or Wikipedia entries where they are warranted.
- Consensus sources: models favor claims that several unrelated sources agree on.
- Third-party citations: answers often cite review platforms, Reddit threads and trade press ahead of vendor pages.
- Expert quotations and proprietary data: original figures give other sites a reason to cite and name you.
- Page-level answer formatting: answer-first paragraphs, comparison tables and clear headings are easier for a model to quote.
- Technical eligibility on Google: noindex tags or snippet restrictions such as nosnippet and max-snippet can exclude pages from AI Overviews. For product brands, Merchant Center feed accuracy matters, including titles, GTINs, availability, price and product highlights that match the product pages.
Match the pattern in your audit to the likely cause:
| Pattern | Likely cause | Business impact | Next action |
|---|---|---|---|
| Absent | Weak entity signals, no third-party coverage | Never considered | Fix entity data and earn review and press coverage |
| Named but not described | Thin or inconsistent descriptions across sources | Listed, then ignored | Align positioning across your site, review profiles and directories |
| Described but not recommended | Competitors own the use-case evidence | Lost at shortlist | Publish use-case and comparison pages and seed peer reviews |
| Recommended only in comparison prompts | Missing from the category lists that answers cite | Wins only buyers who already know you | Get into the cited "best of" pages and Reddit threads |
| Recommended in category prompts | Consensus is established | Pipeline from AI answers | Defend accuracy and refresh evidence |
Takeaway: most AI visibility gaps come from missing evidence on third-party sites rather than from formatting on your own site.
AI Visibility Apps and Tools
AI visibility tools automate the prompt runs and the logging, but they differ in coverage, collection method and price. For a broader list, see our review of AI visibility tracking platforms.
| Tool | Primary job | AI surfaces | Collection | Pricing | G2 rating |
|---|---|---|---|---|---|
| Semrush AI Visibility Toolkit | AI visibility tracking inside an SEO suite | ChatGPT, Perplexity, Gemini, Claude, Copilot, AI Overviews, AI Mode | Mainly API, some browser sessions | About $99/month per domain billed annually; 50 extra prompts for $60/month; extra users $45/month | 4.5 (3,945 reviews) |
| Profound | Enterprise AI search visibility tracking | The same seven surfaces | API, browser sessions, panel data | Custom enterprise quote; 7-day trial with 50 prompts per day | 4.6 (1,124 reviews) |
| Brandwatch | Social listening (owned by Cision) | Social, forums and news; not confirmed for AI assistants | Continuous mention tracking | Quote-based; reported from about $800/month to $15,000+/month | 4.4 (709 reviews) |
Semrush sorts prompts into Strong, Weak and Missing and refreshes prompt rankings daily and brand data weekly. Sentiment reporting is inconsistently documented, and AI referral traffic is not tracked directly.
Profound adds AI referral traffic and sentiment, with a weekly index. It also offers SSO, user roles, approval workflows, an audit trail and SOC 2 Type II.
G2 reviews read in October 2026 make a few points relevant here:
- Semrush users call its AI Overview and AEO tracking increasingly valuable.
- Profound reviewers want more explanation of why scores change, not just that they did, plus broader LLM and country coverage. Some note that "Average Position Rank" can surface generic non-brand entries.
- Brandwatch reviewers report inconsistent sentiment accuracy on sarcastic and multilingual posts.
There are open-source options too. Bright Data's GEO/AEO Tracker is a local-first dashboard you can clone from GitHub, and GetCito offers an open-source AI Visibility Checker. A small team should start with these tools or the Semrush toolkit. Peec AI, Otterly.AI, Ahrefs Brand Radar and Scrunch are also worth comparing.
Takeaway: trackers show where you are missing; closing the gap takes separate work.
How Tellr Turns AI Visibility Gaps Into Published Work
Tellr is a managed earned-visibility program that measures where a brand is missing and then produces the work that changes it. Each week, Tellr tracks the category's queries on Google. It records the organic results, the AI Overview and the discussions block, and labels every cited domain as your own site, a competitor, Reddit, social, a review site or a reference, with week-over-week movement. That data decides what the senior team builds next. Tellr does not offer a prompt-by-prompt dashboard of every assistant. It is built for marketing teams spending $10k+ a month at companies worth $500M+ or with 200+ employees, so smaller teams are better served by a self-serve tracker.
- Comparison pages, reviews and answer-shaped articles, written to be quoted by ChatGPT, Perplexity, Gemini, Claude and AI Overviews, and published to WordPress
- Reddit replies, checked against subreddit guidelines and placed behind an approval gate
- Ready-to-run ad creative based on category ad intelligence
- A weekly digest, a monthly program review, API access and "Tellr for Claude" via MCP
How to Report AI Visibility Every Month
Report AI visibility monthly to executives, rerun a core prompt subset more often, and repeat the full audit each quarter. A workable cadence looks like this:
- Weekly: track Google AI Overview citations and 15 to 20 high-intent prompts. Our guide to weekly AI brand monitoring covers the setup.
- Monthly: send an executive report on recommendation rate for non-branded prompts, share of AI voice against three to five competitors, citation share by source type, and the count of factual errors.
- Quarterly: run the full audit of 40 to 60 prompts with reruns, plus an extra run after any major model release.
Leading indicators move before recommendations do. Watch for these signs:
- Depth scores rise
- The brand picks up new mentions in comparison prompts
- Third-party pages that name you start to appear among cited sources
- Branded search lifts
Give each one a full quarter before you judge the work. AI visibility improves for teams that measure the same way every month against the same competitors and put their effort into the sources that answers already trust.
FAQ
What is AI visibility?
AI visibility measures how often AI-generated answers mention, describe, cite or recommend your brand across tools such as ChatGPT, Perplexity, Gemini, Claude, Copilot and Google AI Overviews. The article breaks it into five metrics: mention rate, recommendation rate, citation frequency, sentiment and positioning, and share of AI voice.
How do you run a repeatable AI visibility audit?
Use a fixed set of 40 to 60 prompts across branded, category, comparison and high-intent buyer searches; keep run conditions consistent; rerun each prompt three to five times per engine within the same 48-hour window; and log every answer verbatim, including brands named, citations, errors, model version and date.
Why do AI answers mention some brands and skip others?
According to the article, brands are more likely to appear when models can recognize them as distinct entities and when multiple independent sources describe them consistently. Missing visibility is often caused by weak entity signals, limited third-party coverage, thin descriptions across sources or a lack of evidence on the review sites, Reddit threads and trade press that AI answers already cite.
Which prompts matter most in AI visibility reporting?
Non-branded category prompts and high-intent buyer prompts matter most because that is where AI answers build shortlists. Branded prompts usually show whether the model recognizes your brand, but they do not reveal whether you are being considered when buyers are still comparing options.
How often should AI visibility be reported?
The article recommends tracking Google AI Overview citations and 15 to 20 high-intent prompts weekly, sending a monthly executive report on recommendation rate, share of AI voice, citation share and factual errors, and running a full audit of 40 to 60 prompts each quarter or after a major model release.