The AI Visibility Audit Checklist (2026 Edition)
· 6 min read · By Perciva Team
An AI visibility audit answers one question: when buyers ask AI engines about your category, your product, and your competitors, what do they actually hear — and how much of it is accurate, current, and in your favor? This checklist walks through the full audit in six phases: question mapping, engine coverage, presence, accuracy, citations, and competitive position, with a scoring rubric so the result is a number you can track quarter over quarter instead of a pile of screenshots.
Run it before you invest in any generative engine optimization work. An audit tells you where you're losing; optimization without one is guessing.
Before You Start
- Pick your engines. At minimum: ChatGPT and Gemini. Add Perplexity and Copilot if your buyers skew technical or enterprise. Auditing one engine tells you about that engine, not about your AI visibility.
- Use clean sessions. Logged-in chat history personalizes answers. Audit in fresh sessions so you see what a new buyer sees.
- Decide the unit of record. Store full verbatim answers with dates. Summaries and screenshots make phase 4 (accuracy) and future comparisons much harder.
Phase 1: Map the Questions Buyers Actually Ask
The audit is only as good as its question set. Build 20–40 buyer-intent prompts across five types:
- Category questions: "best [category] software for [segment]"
- Comparison questions: "[You] vs [Rival] — which should I choose?"
- Alternative questions: "alternatives to [Rival]" (and, uncomfortably, "alternatives to [You]")
- Capability questions: "does [You] support [integration / compliance / feature]?"
- Pricing questions: "how much does [You] cost?" and "is [You] worth it?"
Source these from sales calls, support tickets, and community threads rather than inventing them at your desk — phrasing changes answers. Our prompt library has category-by-category starting points.
Phase 2: Engine Coverage
Run every question on every engine you chose. Log, per answer:
- Date, engine, and exact question text
- The full answer, verbatim
- Which products are named, in what order
- Which product (if any) the answer recommends
- Every cited source URL, if the engine shows citations
Phase 3: Presence — Are You in the Room?
For each category and alternative question, score your presence:
- Named first / recommended: the answer leads with you or picks you.
- Named: you appear, but the recommendation goes elsewhere or nowhere.
- Absent: you don't appear at all.
Aggregate this into a simple AI share of voice figure: the percentage of category-level answers that name you, and the percentage that recommend you. These two numbers are the headline of the audit.
Weight matters more than the average suggests: being absent from your single most-asked comparison question is worse than being absent from five long-tail ones. Mark your five highest-stakes questions before you score, and report their results separately from the aggregate — a healthy overall presence number can coexist with a losing record exactly where deals are decided.
Phase 4: Accuracy — Is What They Say True?
Go through every answer that mentions you and extract each factual claim: pricing, features, integrations, compliance, company facts. For each claim, mark it correct, outdated, or wrong. Pay special attention to:
- Pricing: old tiers and retired plans are the most common stale claims.
- Feature negatives: "does not support X" statements — these kill deals silently and are often simply out of date.
- Positioning: descriptions that anchor you to a segment you've outgrown ("a tool for small teams").
Phase 5: Citations — Where Do the Answers Come From?
List every domain cited across your answers and bucket them: your own properties, neutral third parties, and competitor-owned content. Two findings matter most: which non-brand domains the engines trust for your category (those are outreach and placement targets), and whether any answer about you is built on a competitor's comparison page. The follow-up work here is a citation gap analysis — finding the sources AI trusts where you're absent.
Phase 6: Competitive Position
Re-read the comparison and alternative answers from the rival's perspective. Who wins each head-to-head? Which rival appears most often across all category questions? Note the exact language used to frame you against them — those phrases are what buyers repeat in first sales calls.
The Scoring Rubric
| Dimension | What you measure | Score 0–5 means |
| Presence | % of category answers naming you | 0 = absent everywhere, 5 = named in nearly all |
| Recommendation | % of answers picking you as the choice | 0 = never the pick, 5 = the default pick |
| Accuracy | % of extracted claims that are correct and current | 0 = mostly wrong, 5 = fully current |
| Citation ownership | Mix of your/neutral/rival sources behind answers | 0 = rival-led, 5 = you + strong neutrals |
| Competitive framing | How head-to-heads and framing language treat you | 0 = consistently unfavorable, 5 = consistently favorable |
Total the five dimensions for a 0–25 audit score. The absolute number matters less than the movement: re-run the same question set quarterly and track the delta.
How Long Does This Take?
Budget honestly or the audit stalls at phase 2. For a 30-question set on two engines, expect roughly: half a day to run and capture 60 answers by hand (clean sessions, full copy-paste, source logging), half a day for claim extraction and accuracy grading, and a day for citation bucketing, competitive read-through, scoring, and the write-up — call it two to three working days spread across a week. The second audit is meaningfully faster because the question set, counting rules, and rubric already exist; only the answers are new. Automated capture collapses the first half-day to minutes, which is why teams that audit quarterly almost always end up automating phase 2 first.
Common Findings and What They Mean
Most first audits surface one of four recognizable patterns:
- Strong presence, weak recommendation. You're named everywhere and picked nowhere — the "known alternative" trap. The fix is rarely more mentions; it's better comparison content and clearer differentiation on the questions where the pick happens.
- Accuracy problems concentrated in pricing. Almost always stale third-party roundups plus a pricing page engines parse poorly. Fixable in weeks, and usually the highest-ROI item on the list.
- Great on one engine, absent on another. Engines trust different source ecosystems. Your content strategy has been feeding one of them by accident; the citation phase tells you what the other one eats.
- Rival-owned citations under your own comparisons. The engine describes you using your competitor's comparison page. Until you publish a better source for that question, you're letting the rival write your answer.
Audit Output Checklist
- One-line verdict: your recommendation rate and top rival, in plain language
- The 5-dimension scorecard with this quarter's numbers
- Top 5 wrong or outdated claims, each with the page that needs updating
- Top 5 questions where a rival takes the recommendation
- Citation gap list: trusted domains where you have no presence
- Owner and deadline for each fix
From One-Off Audit to Ongoing Program
A single audit is a snapshot; AI answers move with model updates and competitor content, so the snapshot ages fast. The teams that get value from this treat the audit question set as a permanent monitoring panel and re-check it weekly or monthly — see our walkthrough of how to do an AI visibility audit for a condensed version of this process, and our guide to measuring generative engine optimization for what to track once fixes start shipping. Perciva automates phases 2 through 6 on a schedule if you'd rather not run them by hand.