AI Share of Voice: How to Benchmark Against Competitors
· 6 min read · By Perciva Team
AI share of voice (AI SOV) is the percentage of AI answers to your category's buyer questions in which your brand appears — and, in its stricter form, the percentage in which you're the recommended pick. It's the closest thing generative engines have to a rank tracker, and it's the metric that makes "how are we doing in AI answers?" answerable with a number instead of anecdotes.
This guide defines the metrics precisely, then walks through a benchmarking method that survives AI's non-determinism — the property that breaks most naive measurement attempts.
The Metrics, Defined
Teams use "share of voice" loosely, which makes benchmarks incomparable. Fix the definitions first:
| Metric | Definition | What it tells you |
| Mention rate | % of answers (across your question set) that name your brand | Are you in the conversation at all? |
| Recommendation rate | % of answers where you are the explicit pick | Are you winning the conversation? |
| First-mention share | % of list-style answers where you're named first | Position inside answers buyers skim |
| Citation share | % of cited sources that are your domains | How much of the answer's evidence you own |
| Relative SOV | Your mention rate ÷ (sum of mention rates of tracked brands) | Your slice of the category's total AI presence |
Mention rate flatters you; recommendation rate pays the bills. A brand can be mentioned in 80% of answers and recommended in 10% — "in the room" but never the pick. Track both, and lead reports with recommendation rate.
The Benchmarking Method
- Fix the question set. Benchmarks require a stable panel: 20–50 buyer questions covering category picks, comparisons, alternatives, and use-case-specific asks. Changing the questions changes the number, so version the set — a documented prompt pack — and keep it constant within a benchmarking period.
- Fix the engines and modes. ChatGPT with and without browsing produce different answers. Decide which engines and modes are in scope and hold them constant.
- Sample, don't spot-check. The same question to the same engine can produce different answers across runs. A single run per question is a coin flip wearing a lab coat. Run each question multiple times per period (or at minimum across multiple days) and compute rates over all runs.
- Count with consistent rules. Decide up front what counts as a mention (name in the answer body — not only inside a URL), and what counts as a recommendation (an explicit pick or "best for your case is..." — not mere presence in a list). Write the rules down; the person counting in Q4 won't remember Q2's judgment calls.
- Score every tracked competitor the same way. The same runs, the same counting rules, applied to each rival. Your recommendation rate only means something next to theirs.
- Report the spread, not just the average. "Recommended in 40% of runs" reads very differently if it's stable across engines versus 80% on Gemini and 0% on ChatGPT. Per-engine breakdowns tell you where to work.
How many runs is enough? There's no magic number, but the intuition is simple: the closer two brands' rates are, the more runs you need before the gap means anything. Three runs per question per engine per month is a practical minimum for trend reporting; if you're trying to detect a 5-point shift between near-tied rivals, you need substantially more — or you should report the pair as "contested" rather than pretending to resolution the sample can't support. Weight engines by your buyers, too: if your prospects overwhelmingly use ChatGPT, a blended average that dilutes it with engines they don't use is a vanity aggregate.
A Worked Example (Hypothetical)
Say you monitor 20 questions on two engines, four runs each per month: 160 answers. Your brand appears in 96 of them — mention rate 60%. It's the explicit pick in 40 — recommendation rate 25%. Your nearest rival is mentioned in 120 (75%) and picked in 56 (35%). Relative SOV across the three brands you track: you 33%, rival A 41%, rival B 26%.
Now the reading. The headline isn't the 60% — it's the 10-point recommendation gap to rival A. Cut by question type, you find your recommendation rate is 45% on head-to-heads but 8% on open category questions: buyers who already know you win the comparison, but you're losing the "best tool for..." shortlists where rival A's content dominates the citations. That's a precise, workable diagnosis — and none of it was visible in the mention rate. This is the analytical payoff of keeping metrics separated and question-level detail intact.
Setting Targets Without Fooling Yourself
- Baseline first, targets second. Run two full periods before committing to any goal; the first period's number always moves once your sampling stabilizes.
- Target the gap, not the absolute. "Close the recommendation gap to rival A from 10 points to 5 in two quarters" survives model refreshes better than "reach 40% SOV," because refreshes move everyone's absolute numbers at once.
- Pair every SOV target with an accuracy floor. Being recommended more while engines misquote your pricing is winning the wrong game; track both, and treat a false-claim spike as overriding any SOV win.
- Expect step changes. SOV moves in steps (a model update, a big citation shift), not smooth curves. Judge trends across quarters, not weeks.
Benchmarks Worth Having
- You vs. category leader: the gap in recommendation rate on category questions.
- You vs. nearest rival: head-to-head win rate on direct comparison questions.
- You vs. yourself: the trend line — the benchmark that turns SOV into a program metric. Same panel, same method, every period.
- Segment cuts: SOV on enterprise-flavored questions vs. SMB-flavored ones often diverges sharply and reveals positioning problems no aggregate shows.
Pitfalls That Invalidate Benchmarks
- Non-determinism ignored. One run per question produces numbers that swing week to week for no real reason. If your SOV moves 15 points in a week, check sample size before celebrating or panicking.
- Personalized sessions. Logged-in history skews answers. Measure from clean sessions.
- Leading prompts. "Why is [You] the best?" inflates SOV and measures nothing. Questions must be neutral, phrased the way buyers actually phrase them — see how buyers actually ask AI.
- Counting rule drift. If "mentioned" quietly becomes "mentioned or cited," your trend line is fiction.
- Averaging away the story. A stable aggregate can hide a critical question flipping to a rival. Keep question-level detail underneath the headline number — displacement hides in averages (our displacement detection playbook covers catching it).
Cadence and Reporting
Weekly scans feed monitoring; monthly aggregates make good benchmarks; quarterly deltas belong in leadership decks. Present three numbers — mention rate, recommendation rate, and the top rival's recommendation rate — plus the question-level wins and losses behind them. Resist the composite-index temptation: a single blended "AI score" that mixes mentions, recommendations, and citations feels executive-friendly but hides which lever moved, and the first question anyone asks about a composite is what's inside it anyway. For how SOV fits into a broader measurement stack alongside citation and accuracy metrics, see how to measure generative engine optimization.
Running this by hand is doable but tedious: the sampling requirement is what breaks spreadsheets. Perciva runs the panel across engines on schedule and computes these rates per question and per rival — you can see the output format in our sample report.