Citation Gap Analysis, Step by Step
· 6 min read · By Perciva Team
Citation gap analysis is the process of collecting the sources AI engines actually cite when answering your buyers' questions, classifying each source by who controls it, and identifying the questions where none of the cited sources say what you need said. The output is not a report — it is a ranked to-do list: pages to create, pages to fix, and third-party placements to earn.
It is the highest-leverage analysis in GEO because citations are where answers come from. Search-grounded engines like Perplexity, ChatGPT search, and Gemini build answers out of retrieved sources; if every source behind "best [your category] software" was written by your competitor or ignores you, the answer is decided before generation begins. The citation gap is that structural disadvantage, made visible and fixable. Here is the full method.
Step 1: Assemble the Question Panel
Start from questions with money attached, not questions you find flattering. Pull from discovery-call recordings, sales objections, support tickets, and search query data. Cover four types: category questions ("best [category] tools for [ICP]"), comparison questions ("[You] vs. [Rival]"), capability questions ("does [You] support [feature]?"), and trust questions ("is [You] SOC 2 compliant?"). Twenty to forty questions is enough to be representative while staying tractable. Keep the panel fixed — the analysis only compounds if you re-run the same questions later.
Step 2: Collect Answers and Citations
Run every question through the citing engines — Perplexity, ChatGPT with search, and Gemini are the practical set — and record, per run: the full answer text, every cited URL, and which parts of the answer each citation supports. Run each question more than once; retrieval varies between runs, and a source cited in four of five runs is a different fact than one cited once. Store the verbatim answers. They are your evidence base, and you will need them when a fix later changes an answer and you want proof.
Step 3: Classify Every Cited Source
Tag each unique cited URL into one of four buckets:
- Owned — your domain: docs, pricing, blog, comparison pages.
- Earnable — third-party sources you could plausibly influence: review platforms, industry listicles, community threads, partner content.
- Competitor-owned — a rival's domain, including their comparison pages about you.
- Uninfluenceable — encyclopedic or news sources where placement is not realistically actionable.
This classification is where the analysis becomes strategic: the earnable bucket is your PR roadmap, and the competitor-owned bucket is your risk register — every answer grounded in a rival's "Us vs. You" page is an answer written by your competitor's marketing team.
Step 4: Build the Gap Matrix
Lay questions against citation buckets and count. Three patterns demand action, in order of severity:
- Zero-owned questions — engines answer entirely from sources you do not control. If the question is commercial ("pricing," "vs."), this is urgent.
- Rival-grounded questions — competitor-owned sources dominate the citations. Expect the answer's framing to match.
- Thin-answer questions — few citations of any kind, meaning engines lack good sources. These are open ground: the first strong page often becomes the canonical citation.
Step 5: Diagnose Each Gap
For every question where you are absent from citations, there is a specific reason. Work through them in order:
- No page exists. You never wrote the page that answers this question. Most common, most fixable.
- The page is unreachable. Blocked by robots.txt or your CDN's bot rules, gated, or JavaScript-only. Check crawler access before rewriting anything.
- The page is unquotable. It exists and is crawlable but buries the answer, hedges, or lacks extractable structure — headings, direct answers, lists.
- The page loses on authority. It is fine, but engines prefer a higher-trust third party for this question type. Engines systematically prefer independent sources for "best" and "vs." questions — a bias explained in our source authority entry. The fix is earned placement, not another owned page.
Step 6: Prioritize by Answers Influenced
Not all gaps are equal. Rank fixes by: commercial weight of the question (comparison and pricing beat informational), number of answers the source influences (one listicle cited across six questions outranks six single-question fixes), and feasibility. A useful heuristic: fix owned-content gaps first (fully in your control, days not months), then pursue the top three earnable sources by influence — for those, the playbook is digital PR for AI citations.
Step 7: Act, Then Re-Measure
Ship the fixes, then re-run the identical panel after a few weeks and diff: did your owned-citation count rise, did any rival-grounded question flip, did new sources appear? Citation sets shift with model and index updates even when you do nothing, which is why one-off analysis decays — the teams that win treat this as a loop, not a project. Continuous citation monitoring is the difference between knowing your gaps once and knowing them now.
A Worked Example
A fictional but representative run: a 30-question panel for a mid-market HR SaaS, executed across Perplexity, ChatGPT search, and Gemini, three runs per question. The harvest yields 214 unique cited URLs collapsing to 41 domains. The matrix shows: 9 questions with zero owned citations, 6 of them commercial; two industry listicles ("Top HR Platforms 2026" on two trade blogs) cited across 11 different questions; the main rival's "vs." page grounding 4 of 5 comparison questions; and G2 present on every "best" and "alternatives" question, quoting a three-year-old profile description.
The resulting priority list writes itself: fix the G2 description this week (one hour, influences a dozen answers); build the two missing comparison pages (in your control, counters the rival's framing); pitch inclusion in both listicles (two emails, 11 answers of leverage); and add a pricing FAQ, because the pricing question showed engines guessing from a stale third-party article. Note what did not make the list: the encyclopedic citations (uninfluenceable) and the informational questions where owned content already appears. The matrix's job is exactly this — separating the four fixes that move answers from the forty that merely feel productive.
Frequently Asked Questions
How often should the analysis be re-run?
Quarterly as a full exercise, with the caveat that citation sets drift continuously — engines re-crawl, listicles update, models refresh. If a quarter is your cadence, accept that you are sampling a moving target; if the category is competitive enough that answer flips cost real pipeline, standing monitoring replaces the quarterly ritual entirely.
Which engines should be in scope?
The ones that cite: Perplexity, ChatGPT with search enabled, and Gemini give you three materially different retrieval systems. Non-citing chat modes still matter for perception, but they cannot feed a citation analysis — their influence shows up in the answer text instead, which is a claims-accuracy question rather than a gap question.
What Teams Usually Find
Running this for the first time reliably surfaces the same handful of surprises: a competitor's comparison page quietly grounding half your "vs." answers; review platforms cited far more than your own site on category questions; documentation outperforming marketing pages for capability questions; and at least one high-intent question where every engine is guessing from thin sources. Each of those is an action, and none of them is visible from inside your analytics stack — which is precisely the point of doing the analysis. Perciva automates this loop end to end, from panel runs to citation extraction to the ranked gap list, if you would rather act on it than assemble it.