AEO tracking is the repeatable measurement of whether AI answer engines such as ChatGPT, Perplexity, Gemini, Claude, Copilot and Google AI Overviews name, cite or recommend your brand when buyers ask a fixed set of category questions. It runs on a schedule and is benchmarked against competitors. Done well, it shows which prompts you win, which engines ignore you and which third-party sources the models trust instead of your site. In October 2026, most tools still lead with a single visibility score, often built on a handful of cherry-picked prompts. This guide starts with the measurement itself. It covers what counts as a mention, which metrics matter, how to sample prompts, how to test a tracker before you buy one, and how to turn AEO tracking into a weekly report your team can act on.
Key takeaways
- A brand can be cited in an AI answer without being named, and named without being recommended, so each signal needs its own metric.
- Answer mention rate, citation rate and share of voice only mean something when you calculate them on a fixed, segmented prompt set that is rerun several times per period.
- AI trackers collect answers through API calls, browser sessions or panel data, and the collection method affects how far you can trust the numbers.
- Validate every tracker by hand on a small prompt sample before purchase, checking repeatability, entity matching and freshness lag.
- Tracking pays off only when each pattern in the data maps to a specific fix, such as a comparison page, a review-site presence or a Reddit reply.
What counts as a mention in an AI answer
A mention is any appearance of your brand in an AI-generated answer, and it splits into signals of very different commercial value. The practice of improving those signals is called generative engine optimization, also called answer engine optimization. Measuring it starts with a four-step ladder:
- Indexed: the engine can retrieve your pages. This is necessary, but the buyer never sees it.
- Cited: one of your URLs appears as a source behind the answer.
- Mentioned: your brand name appears in the answer text.
- Recommended: the answer tells the buyer to consider or choose you.
These levels do not stack neatly. A security vendor's blog post can be the cited source for a definition while the answer recommends three competitors. A tracker that counts both as "visibility" will flatter you.
| Mention type | What it looks like | How to record it |
|---|---|---|
| Explicit brand mention | Brand named in the answer text | Yes/no per answer, plus position (first, middle, last) |
| Linked citation | Your URL shown as a clickable source | URL, page type, citation position |
| Unlinked citation | Your data or wording used, with the brand credited but not linked | Flag manually; most tools miss it |
| Recommendation without brand mention | Your page is the source, but the answer names no vendor | Count as a citation, not a mention |
| Competitor mention | Rival brands named in the same answer | All brands named, in order |
| Framing and sentiment | "Best for enterprises" vs "expensive and dated" | Positive / neutral / negative, plus the descriptor used |
Entity ambiguity and separate scopes
Brands with common names need disambiguation rules before any counting starts. The acronym itself shows the problem. AEO is also the stock ticker of American Eagle Outfitters, so a naive string match on "AEO" mixes retail news into marketing data. Write match rules that combine product names, domains and category context, and keep an exclusion list. Then track three scopes separately, because they move independently:
- Category: do you appear when buyers ask for "best cloud security platforms"?
- Product: do individual products appear for their use-case prompts?
- Executive: are your leaders named as experts on category topics?
The metrics that make AEO tracking measurable
AEO tracking becomes measurable when you express each signal as a rate over a fixed prompt set. A one-off screenshot does not count. These are the metrics worth reporting:
| Metric | Formula or definition |
|---|---|
| Answer mention rate | Answers naming your brand ÷ total answers collected |
| Citation rate | Answers citing at least one of your URLs ÷ total answers |
| Share of voice | Your brand mentions ÷ all tracked brand mentions across the prompt set |
| Engine-specific visibility | Mention rate calculated per engine, never reported only as a blended figure |
| Prompt category coverage | Prompt clusters where you appear at least once ÷ total clusters |
| Average placement | Mean position of your brand among the brands named in the answer |
| Sentiment distribution | Percentage of mentions that are positive, neutral and negative |
| Mention persistence | Share of prompts where you appeared last period and still appear this period |
| Competitor displacement rate | Prompts where you replaced a rival ÷ prompts where you were absent last period |
Our deeper guide on how often AI answers mention your brand covers how to benchmark these rates. The key discipline is the denominator. If you change the prompt set, every trend line breaks.
Worked example: a 100-prompt set
Take an illustrative set of 100 prompts run once each on four engines, which gives 400 answers. Your brand is named in 92 of them, so the answer mention rate is 92 ÷ 400 = 23%. Your URLs are cited in 60 answers, a 15% citation rate. The 400 answers contain 460 brand mentions across you and your tracked competitors, so your share of voice is 92 ÷ 460 = 20%.
Now segment the same data. The 40 informational prompts produce 160 answers, and you appear in 64 of them (40%). The 60 commercial prompts produce 240 answers, and you appear in 28 (11.7%). The blended 23% hides the real finding. Models use you to explain the category but do not recommend you when buyers are choosing.
Leading and lagging indicators
- Leading: new citations of your pages, third-party review and Reddit coverage, citation share on Google AI Overviews, and mention persistence. These move within weeks after content ships.
- Lagging: commercial-prompt mention rate, recommendation share, AI referral traffic, and pipeline influenced by AI-sourced visits. These confirm whether the movement reached buying decisions, and they usually trail by a quarter.
How to build and sample a representative prompt set
A representative prompt set mirrors how your buyers ask questions across the funnel, and it is segmented so you can measure each cluster on its own. Build it from search queries, sales call notes, support tickets and the Reddit threads your buyers read. Then tag every prompt on these dimensions:
| Dimension | Example prompt |
|---|---|
| Branded vs non-branded | "Is [brand] good for SOC 2?" vs "best SOC 2 automation tools" |
| Commercial vs informational | Buying questions vs explanations of the category |
| Funnel stage | Top-funnel problem framing vs bottom-funnel vendor selection |
| Feature and use case | "Tool for CSPM in multi-cloud AWS and Azure" |
| Comparison | "[brand] vs [competitor]", "alternatives to [competitor]" |
| "Best X" | "Best endpoint security for mid-market companies" |
| Troubleshooting | "Why is my cloud bill spiking after enabling logging" |
| Industry and location | "Best password manager for UK healthcare" |
Sampling rules
- Use at least 25 prompts per cluster as a working rule. With that many, one answer moves the cluster rate by four points rather than twenty.
- Run each prompt three to five times per period, because models regenerate answers differently each time. Record a mention when it appears in the majority of runs.
- Check weekly for trend reporting. Daily checks add noise unless you are monitoring a launch or a crisis.
- Fix and log the run conditions: country, language, device, logged-in or logged-out state, and model version.
- Freeze the prompt set for a quarter and add new prompts as a separate cohort, so trends stay comparable.
- Have someone outside the SEO team approve the set. This stops the team from choosing prompts it already wins.
How engines differ
| Engine | Answer and citation behavior | What it means for tracking |
|---|---|---|
| Google AI Overviews | Generated on the results page, with linked sources drawn from Google's index | Classic SEO strength carries over; track cited domains per query |
| Perplexity | Searches the live web for most queries and shows numbered sources | Citation tracking is reliable, and fresh content can appear quickly |
| ChatGPT | Answers from training data unless search is triggered; sources appear when it searches | Results split between "remembered" and "retrieved" answers |
| Gemini | Can ground answers in Google Search results | Overlaps with Google visibility, but not exactly |
| Claude | Uses web search when enabled, otherwise training data | Mentions often reflect long-standing reputation |
| Copilot | Grounded in Bing results, with linked citations | Bing indexation matters more than most teams assume |
AEO tracking tools compared on measurement quality
Judge an AEO tracking tool by what it measures and how it collects answers. The length of its feature list tells you little. API calls are repeatable, but their results can differ from what a logged-in user sees. Browser sessions are closer to real answers but harder to scale. Panel data reflects real user exposure, but it is modeled rather than observed per prompt. For a broader catalogue, see our guide to what each AEO tool measures.
| Tool | Model | Surfaces tracked | Collection method | Refresh | Pricing |
|---|---|---|---|---|---|
| Tellr | Managed program | Google organic, AI Overviews, discussions block | Own search-data pipeline, not a panel | Weekly, with a monthly review | Per engagement, scoped on a call |
| Semrush AI Visibility Toolkit | Self-serve | ChatGPT, Perplexity, Gemini, Claude, Copilot, AI Overviews, AI Mode | Mainly API, some browser sessions | Daily prompt rankings, weekly brand data | About $99/month per domain |
| Profound | Self-serve, enterprise tiers | Same seven surfaces | API, browser sessions, panel data | Core index weekly | Not published in our sources |
| Scrunch AI | Self-serve, enterprise support | Same seven surfaces | API, browser sessions, panel data | Described as continuous; some reviewers report weekly | Not published in our sources |
Tellr
What it measures: Tellr tracks weekly citation share for the category's queries on Google. It records which domains the organic results, the AI Overview and the discussions block cite, labels each as your site, a competitor, Reddit, social, review sites or references, and shows week-over-week movement. It also reports the Reddit threads found, the replies placed and whether each reply is still live, plus where shipped articles rank and get cited.
How it collects: Tellr uses its own search-data pipeline. The measurement steers content that is written to be quoted by ChatGPT, Perplexity, Gemini, Claude and AI Overviews.
Strengths
- Connects each citation gap to a specific fix rather than a score.
- Is governed by a brand brief, claim guardrails, an approval gate and an audit trail.
- Publishes to WordPress, offers API access and provides a "Tellr for Claude" MCP server.
- Runs without GA4 or Search Console access.
Best fit: enterprise marketing teams spending $10k+ a month that want someone to close the visibility gap as well as report it.
Semrush AI Visibility Toolkit
What it measures: brand mentions, citations, share of voice and prompt position. Our sources disagree on whether it reports sentiment, and it does not track AI referral traffic directly. Pricing: about $99/month per domain, billed annually. Extra users cost $45/month each, and each additional 50 prompts cost $60/month. No free trial is stated for the toolkit.
Strengths
- Sits beside the keyword, audit and rank data that teams already use.
- G2 reviewers rate Semrush 4.5/5 across 3,945 reviews as of October 2026, and several call its AI Overview and AEO tracking increasingly valuable.
Limitations
- Reviewers cite limited regional and language coverage.
- Some reviewers question its accuracy on newer AI features.
- Caps on prompts, exports and domains push costs up.
Who should not buy: global brands that need multilingual tracking at scale.
Profound
What it measures: mentions, citation share, share of voice, sentiment, position and AI referral traffic, plus AI crawler activity.
Strengths
- Enterprise governance includes SSO, user roles, approval workflows, an audit trail, multi-brand workspaces and SOC 2 Type II.
- G2 reviewers rate it 4.6/5 across 1,124 reviews as of October 2026, and they single out prompt tracking, citation and source analysis, and competitor benchmarking.
Limitations
- Reviewers report confusing "Average Position Rank" results that include generic non-brand entries.
- Export options are limited, and lower tiers restrict history and API access.
- Some Reddit users question whether its scraping-based approach is reliable.
Who should not buy: small teams without an analyst to turn the data into actions.
Scrunch AI
What it measures: mentions, citations, citation share by domain, share of voice, sentiment, in-answer position, and AI referral and bot traffic. Sitecore acquired Scrunch in June 2026.
Strengths
- Supports custom prompt tracking and competitor comparison.
- Links agent traffic to GA4.
- G2 reviewers rate it 4.6/5 across 73 reviews as of October 2026.
Limitations
- Reviewers mention inconsistent brand mention counts between views.
- Historical and trend data are limited.
- It has no built-in PDF reporting, and personas cannot be edited after creation.
- Prompt credits are consumed per engine.
Who should not buy: teams that need long historical baselines for board reporting.
How to test a tracker, and what no tracker can measure
You test a tracker by reproducing its numbers by hand on a small prompt sample, and you should accept from the start that some AI exposure stays invisible to every tool. Run this test during any trial:
- Pick 20 prompts across your clusters. Run each one manually five times on two engines, logged out, in the target country.
- Compare your manual mention rate with the tool's figure for the same prompts and dates. A gap of more than ten points needs an explanation.
- Rerun the same prompts in the tool a week later, with no content changes, to measure variance.
- Check that the cited URLs in the tool match the sources shown in the live answer.
- Search the tool's data for false positives such as name collisions, partner mentions and former employees.
- Measure the freshness lag between a change in a live answer and its appearance in the data.
Buyer's checklist: score each vendor on engine coverage, prompt capacity, historical retention, citation extraction quality, segmentation, API and BI export, alerting, team permissions, entity disambiguation, international coverage and data retention terms. Ask your compliance team to review how answers are collected before you sign.
What can't trackers measure yet? Trackers cannot see personalized answers shaped by a user's memory and history, hidden system prompts, model updates that change behavior overnight, variance between answer regenerations, or citations shown to some users but not others. Treat every metric as a sampled estimate of exposure rather than a census.
Where Tellr fits in an AEO program
Tellr fits after a tracker has shown you where you are missing. At that point, Tellr's senior team runs one governed program on Tellr's platform to change those answers. Tellr is overkill for startups and small teams, which are better served by a self-serve tracker. Week one produces a category map of the threads, queries and AI answers that matter. Week two agrees the brief and guardrails. After that, the work runs weekly, reported in a digest with a monthly program review. The program delivers:
- Comparison pages, reviews and answer-shaped articles, published straight to your CMS.
- Subreddit mapping, a daily thread radar and guideline-checked Reddit replies behind an approval gate.
- Category ad intelligence and ready-to-run creative.
- Read-only viewer roles for stakeholders outside the marketing team.
Turning tracking into a weekly report and action
A useful AEO report shows the same columns every week and ties each change to an action. You can replicate this scorecard in a spreadsheet or BI tool:
| Column | Example value |
|---|---|
| Prompt ID / cluster / funnel stage | P-047 / comparison / bottom |
| Engine / country / login state / run date | Perplexity / US / logged out / 2026-10-12 |
| Brand mentioned (majority of runs) | Yes, 3 of 5 |
| Placement / sentiment / descriptor | 2nd of 4 / positive / "strong for multi-cloud" |
| Our URLs cited / third-party sources cited | 1 / G2, Reddit, analyst blog |
| Competitors named / change vs last week | 3 / gained vs Competitor B |
For the weekly rhythm of reviewing this data, see our guide to tracking what the models say about you. Then map each pattern to a fix:
| Pattern in the data | Fix |
|---|---|
| Cited but not named | Rewrite the source pages so the brand, product and claim sit in the same sentence |
| Competitors dominate citations | Publish comparison and "alternatives" pages that answer the exact prompt |
| Visible in informational prompts, absent in commercial ones | Build bottom-funnel pages and secure review-site coverage |
| One engine names you, another does not | Check that engine's retrieval source, such as Bing indexation for Copilot |
| Citations come from reviews and Reddit | Invest in those surfaces directly, because the models already trust them |
Choose your setup by maturity. A small team can run manual checks on 50 prompts in a spreadsheet. A mid-sized team with an analyst gets real value from a self-serve tracker. An enterprise with $10k+ in monthly spend and multiple products needs AEO tracking connected to a program that ships the pages, replies and creative that move the numbers.
FAQ
What counts as a mention in AEO tracking?
A mention is any appearance of your brand in an AI-generated answer, but the article separates four different signals: indexed, cited, mentioned and recommended. These do not always overlap, so citation rate, mention rate and recommendation share should be tracked separately.
Which metrics matter most for AEO tracking?
The core metrics are answer mention rate, citation rate, share of voice, engine-specific visibility, prompt category coverage, average placement, sentiment distribution, mention persistence and competitor displacement rate. The article stresses that each metric only means something when it is calculated on a fixed prompt set.
How should you sample prompts for reliable AEO measurement?
Use a representative prompt set built from real buyer questions, then segment it by factors like branded vs non-branded, commercial vs informational, funnel stage, use case and comparisons. As a working rule, the guide recommends at least 25 prompts per cluster and running each prompt three to five times per period, then counting a mention only if it appears in the majority of runs.
How do you test whether an AEO tracker is trustworthy?
The article recommends validating every tracker by hand before you buy it. Pick 20 prompts across your clusters, run each manually five times on two engines, compare your manual mention rate with the tool's numbers, rerun a week later to measure variance, verify cited URLs, check for false positives and measure freshness lag.
What can’t AEO tracking tools measure yet?
Trackers still cannot fully measure personalized answers shaped by user history, hidden system prompts, sudden model updates, regeneration variance or citations shown to some users but not others. The guide advises treating every AEO metric as a sampled estimate of exposure, not a complete census.