To run an AI visibility audit in a week, fix a set of buyer prompts, run them across ChatGPT, Gemini, Perplexity, Claude and Google AI Overviews under controlled conditions, score every answer with one rubric, benchmark competitors, trace the cited sources, and turn the gaps into a prioritized fix plan by Day 7. The audit shows how often, how prominently and how accurately AI engines present your brand when buyers ask category questions. The sprint below covers prompt frameworks, bias controls, a scoring worksheet and a remediation matrix, and it reflects how the engines behave in October 2026.
Key takeaways
- A one-week AI visibility audit needs a fixed prompt set, consistent test conditions and one scoring rubric applied to every answer.
- Unbranded, decision-stage prompts reveal real visibility, while branded prompts mainly test accuracy.
- Citation analysis shows which third-party domains shape AI answers in your category, and those domains usually decide the fix plan.
- Every audit finding should map to a specific fix category, ranked by business impact and implementation difficulty.
What an AI visibility audit measures
The audit records whether AI engines mention your brand for category questions, where they place it in the answer, what they say about it, and which sources they cite. A spot check with an AI visibility checker gives you one snapshot. The audit uses a structured prompt set, repeated runs and competitor benchmarks.
| An AI visibility audit is | An AI visibility audit is not |
|---|---|
| A fixed, reproducible prompt set run across several engines | A handful of ad hoc questions typed into ChatGPT |
| A scored dataset covering presence, position, sentiment and citations | A yes/no check on whether the brand appears |
| A competitor benchmark on the same prompts | A view of your brand in isolation |
| A source map of the domains that feed the answers | A rank report for blue links |
| A prioritized fix plan with a re-audit date | A one-off screenshot deck |
The core signals are share of voice, citations, sentiment and citation-weighted impact, which is prominence weighted by how well each mention is sourced. Accuracy matters as much as any of them, because a misstated price or feature can cost more than an absence.
Days 1–2: Set scope and build the prompt set
By the end of Day 2 you should have fixed audit boundaries and a frozen prompt set that every engine receives word for word.
Day 1 scoping checklist
- Markets and locales, such as US English and UK English as separate runs
- Buyer personas and the categories they search, such as "cloud security posture management"
- Brand entities: company name, product names, common misspellings, named experts
- Competitors: 1–3 direct rivals plus 1–2 adjacent alternatives buyers also consider
- Engines and modes: ChatGPT (with and without search), Gemini, Perplexity, Claude, Microsoft Copilot, Google AI Overviews
- Owners: one person runs prompts, one verifies scores, one owns the fix plan
Report direct competitors (same buyer, same budget line) and adjacent ones (a platform suite or open-source option that solves part of the problem) in separate columns. Losing to an adjacent suite calls for different fixes than losing to a direct rival.
Day 2: prompt frameworks by funnel stage
Rewrite your top transactional keywords as the conversational, decision-focused questions buyers type into chat. Aim for 20–30 high-intent unbranded prompts plus a smaller branded set. Avoid two-word prompts, and avoid prompts so branded they force a mention.
| Funnel stage | Prompt template | Example (cloud security) |
|---|---|---|
| Discovery | What are the best [category] for [persona/company size]? | What are the best CSPM tools for a 2,000-person SaaS company? |
| Problem | How do I solve [pain] without [constraint]? | How do I find misconfigured S3 buckets across 40 AWS accounts without adding headcount? |
| Comparison | [Competitor A] vs [Competitor B] for [use case]? | Wiz vs Orca for multi-cloud compliance reporting? |
| Validation | Is [brand] good for [use case]? What do users complain about? | Is [your brand] reliable for SOC 2 evidence collection? |
| Branded | What does [brand] cost / integrate with / do? | Does [your brand] integrate with Jira and Slack? |
Unbranded prompts measure visibility. Branded and validation prompts measure accuracy and sentiment. Score the two groups separately so a strong branded result does not hide weak category presence.
Follow-up prompts and normalization
- Write one follow-up per discovery prompt, such as "Which of those is best for a regulated fintech?", to test whether your brand survives narrowing.
- Freeze the exact wording in a sheet with a prompt ID, and never paraphrase between engines.
- Add a fixed context line (company size, region) instead of letting each tester improvise.
- Ask engines to "list sources" only if every engine gets the same instruction.
Pull prompt wording from sales call notes, Reddit threads and support tickets. Buyers phrase questions differently from the way your keyword tool does.
Day 3: Collect answers under controlled conditions
On Day 3, run every prompt across every engine with the same session settings, repeat each unbranded prompt, and keep the full evidence for every run.
Bias controls
- Session hygiene: a fresh chat per prompt, with memory and custom instructions off and chat history cleared
- Account state: clean test accounts, not employee accounts that have discussed your brand
- Geolocation: a fixed VPN exit or a dedicated browser profile per locale, logged for every run
- Browsing mode: record whether web search was on, since ChatGPT with and without search behave like different engines
- Personalization: run Google AI Overviews logged out in a private window to reduce history effects
- Timing: spread repeat runs across morning and afternoon to catch variation
Sample size
Run each unbranded prompt at least three times per engine, and count a brand as reliably present only when it appears in two of three runs. For example, 25 prompts × 4 engines × 3 runs gives 300 answers, enough to compare yourself with 3–4 competitors without single runs skewing the result. With fewer than roughly 100 answers, report directions rather than precise percentages.
Evidence to capture per run
- Prompt ID, engine, model or mode, locale, date and time
- Full answer text in the log, plus a full-page screenshot
- Every cited URL, in order
- Every brand named, in order of appearance
- Attributes stated about your brand, such as price, features and integrations
Ongoing AEO tracking of AI answers depends on the same discipline. If a run cannot be reproduced from the log, it cannot be compared next quarter.
Days 4–5: Score answers and benchmark competitors
On Days 4 and 5, apply one rubric to every logged answer, check claims against the facts, and compare your weighted impact with competitors on the same answers.
Day 4: scoring worksheet
| Signal | Rule | Points |
|---|---|---|
| Prominence | Primary recommendation / secondary mention / brief reference / absent | 5 / 3 / 1 / 0 |
| Sentiment | Positive / neutral / negative | +1 / 0 / −1 |
| Citation support | Mention backed by a cited source that names you | +1 |
| Accuracy | Any false claim about your brand | Answer scores 0, logged as a fix |
AI Visibility Score = (sum of points ÷ maximum possible points) × 100, with a maximum of 7 points per answer (5 + 1 + 1). For example, 300 answers give 2,100 possible points, so a brand earning 630 points scores 30.
Edge cases and a worked example
- Uncited answers: score prominence and sentiment, give no citation point, and flag them as "unsupported", since they shift more between runs
- Hallucinated claims (wrong pricing, discontinued products, invented features): set the answer to 0 and record the exact false statement
- Ambiguous mentions (a parent company or similarly named product): score 0 and file under entity clarity
For example, a high-scoring answer to "best CSPM tools for a 2,000-person SaaS company" names your brand first, describes multi-account coverage accurately and cites an independent review page, which scores 5 + 1 + 1 = 7. A low-scoring answer lists you fifth, says you "lack Azure support" when you have it, and cites nothing. It scores 0 and goes on the fix list.
As working bands, 0–20 means largely absent, 21–40 present but secondary, 41–60 competitive, and above 60 category-leading on your prompt set. Calibrate them against competitors rather than treating them as absolute. The explainer on what a good AI visibility score looks like covers this in depth.
Day 5: share of voice versus weighted impact
Raw mention counts treat a fifth-place footnote the same as a top recommendation, so weight them. Weighted impact = Σ (prominence points × citation multiplier × engine weight), with a multiplier of 1.5 for source-backed mentions and 1.0 otherwise. Engine weights should reflect where your buyers research. One example split is ChatGPT 0.35, Google AI Overviews 0.30, Perplexity 0.20 and Gemini 0.15.
| Brand (example) | Mentions (300 answers) | Mention share of voice | Weighted impact | Weighted share |
|---|---|---|---|---|
| Your brand | 96 | 24% | 118 | 19% |
| Competitor A (direct) | 132 | 33% | 246 | 40% |
| Competitor B (direct) | 104 | 26% | 150 | 25% |
| Competitor C (adjacent) | 68 | 17% | 98 | 16% |
In this illustrative data, your mention share (24%) looks close to Competitor B's, but your weighted share (19%) is lower because your mentions sit lower in answers and are rarely cited. That gap points to source authority, not awareness. Segment the table by funnel stage too, because brands often appear in discovery prompts and vanish in comparison prompts.
Day 5: source analysis
Classify every cited URL as owned content, third-party reviews (G2, Gartner Peer Insights, Capterra), editorial or media coverage, community forums (Reddit, Stack Overflow) or competitor content. Then answer five questions:
- Which 10 domains are cited most across all engines, and are you present on each?
- What share of citations is first-party versus third-party?
- Do aggregator "top 10" listicles get cited over vendor docs or analyst coverage?
- Do cited pages rank beyond page one of Google? If so, a page does not need a high Google rank to get cited, and a well-structured page can earn citations directly.
- Which competitor pages are cited for comparison prompts you have no page for?
Days 6–7: Build the fix plan and report
On Days 6 and 7, turn findings into fixes tied to specific triggers, rank them, and deliver a report stakeholders can act on.
Remediation playbook
| Audit finding (trigger) | Fix category | Typical action |
|---|---|---|
| Ambiguous or confused mentions | Entity clarity | Consistent naming and a clear "what we do" statement on the homepage and About page |
| Absent on problem-stage prompts | On-site content | Answer-shaped articles that state the problem and solution in the first lines |
| Competitor pages cited for "vs" prompts | Comparison pages | Fair, specific comparison pages for each direct competitor |
| Wrong pricing or features | Schema implementation | Organization, Product and FAQPage structured data that matches visible page content |
| Uncited or low-prominence mentions | Citation earning | Placement on the top cited listicles and editorial domains from Day 5 |
| Negative sentiment on validation prompts | Review ecosystem | Fresh reviews on G2 and Gartner Peer Insights, responses to recurring complaints |
| Thin trust signals | Expert authorship | Named authors with credentials, bylined research |
| Conflicting facts across sources | Profile consistency | Aligned descriptions on LinkedIn, Crunchbase, Wikipedia-eligible sources and review profiles |
Prioritization matrix
| Issue | Business impact | Difficulty | Priority |
|---|---|---|---|
| Hallucinated pricing or features | High | Low | Quick win |
| Missing comparison pages | High | Medium | Mid-term |
| Absent from top cited third-party domains | High | High | Strategic |
| Inconsistent profiles | Medium | Low | Quick win |
| Weak authorship signals | Medium | Medium | Mid-term |
Within each band, rank by breadth of exposure (how many prompts the issue affects) and confidence (whether the finding held across repeat runs).
Day 7 report components
- One-page executive summary: score, weighted share versus competitors, top three fixes
- Visibility snapshot by engine and funnel stage
- Evidence bank of response logs and screenshots, plus a citation map of the top cited domains
- Prioritized fix plan with owners, dates and a re-audit date on the frozen prompt set
Close with a 30-minute stakeholder walkthrough covering what you found, why it matters for pipeline, and what to fix first.
Where Tellr fits after the audit
Tellr runs the work an audit points to as one governed program for enterprise marketing teams. Its answer visibility product tracks, every week, who Google and its AI Overviews cite for your category's queries, which keeps the Day 5 source map current. It is built for teams spending $10k+ a month at companies worth $500M+ or with 200+ employees, so smaller teams will usually get more value from a self-serve tracker and in-house fixes. The fixes run inside the same program:
- Comparison pages, reviews and answer-shaped articles built to be quoted by ChatGPT, Perplexity and Google AI Overviews, published to your CMS
- Subreddit mapping, a daily thread radar and guideline-checked Reddit replies behind an approval gate
- Category ad intelligence and ready-to-run creative
- Guardrails, approvals and an audit trail across all of it
Keep the audit running after week one
A one-week audit pays off when you re-run the same frozen prompts on the same engines monthly or quarterly and compare the trend lines.
- Keep prompt IDs, locales, session settings and engine weights unchanged between runs.
- Log model or mode changes, because engine updates can move scores with no change on your side.
- Re-score with the identical rubric and compare weighted share per competitor. A rising share for Competitor A usually traces to a new review, listicle or Reddit thread in the source map.
- Check whether each shipped fix moved the specific prompts it targeted.
- Add new prompts as a separate set so the baseline stays comparable.
Run your first AI visibility audit as a pilot on one category and one locale, then expand the prompt set once the scoring and evidence process holds up under a second run.
FAQ
What does an AI visibility audit actually measure?
It measures whether AI engines mention your brand for category questions, where they place it in the answer, what they say about it, and which sources they cite. The core signals are share of voice, citations, sentiment, accuracy, and citation-weighted impact.
How many prompts and test runs do you need for a one-week audit?
The article recommends 20–30 high-intent unbranded prompts plus a smaller branded set. Each unbranded prompt should be run at least three times per engine, and a brand counts as reliably present only if it appears in two of those three runs.
Why should unbranded and branded prompts be scored separately?
Unbranded prompts test real category visibility, while branded and validation prompts mostly test accuracy and sentiment. Keeping them separate prevents strong branded performance from masking weak visibility on decision-stage category queries.
How is the AI Visibility Score calculated?
Each answer is scored for prominence, sentiment, citation support, and accuracy. The maximum is 7 points per answer: 5 for prominence, 1 for positive sentiment, and 1 for citation support. If an answer contains a false claim about your brand, it scores 0 and is logged as a fix item. The final score is (sum of points ÷ maximum possible points) × 100.
What should the final fix plan include after the audit?
The fix plan should map each finding to a specific action, such as improving entity clarity, creating comparison pages, adding structured data, earning citations on key third-party domains, strengthening review profiles, or aligning inconsistent brand facts across sources. It should also rank issues by business impact, implementation difficulty, breadth of exposure, and confidence from repeat runs.