AI Visibility · Method
AI Brand Visibility Gap Analysis That Survives Noise
Before you compare yourself to a competitor in AI answers, one finding has to be dealt with: ask the same model the same buying question twice and you may get two different winners. Here is how to run gap analysis anyway.
Key takeaways
- One team asked an identical buying question repeatedly, seconds apart, and got a different top brand 4 times in 10. In contested categories the top recommendation was close to a coin toss.
- That makes single-run gap analysis close to worthless. Measure mention frequency across repeated runs of a fixed prompt set, not position in one answer.
- There are two gaps. The mention gap tells you where you are absent. The citation gap tells you which sources produced the answer, and only the second one is actionable.
- Build the prompt set from Search Console rather than imagination, because a self-invented prompt list quietly encodes your own assumptions about how buyers ask.
- Verify AI crawlers can reach your site before anything else. Security and bot-fighting settings commonly block them, and an uncrawlable brand cannot be cited.
01The volatility problem, measured
Start with the experiment that should govern how you read every AI visibility report you will ever see. A team ran a simple test: the same buying-intent question, the same AI, asked repeatedly seconds apart, with no prompt changes.
“It named a different top brand 4 out of 10 times. The full list of recommended brands only overlapped about half the time from one ask to the next. Even after dialing down the randomness settings as much as possible, it still switched its top pick 1 in 4 times.”
Their one exception is as important as the finding: markets with a clear dominant player were stable. In any contested category, the top recommendation was, in their words, basically a coin toss.
That is a single team’s test rather than a controlled study, and we have not replicated it. But it is consistent with what practitioners report elsewhere, including the observation that the same prompt can return completely different recommendations each time with no reliable way to prevent it, which we covered in our GEO versus SEO analysis.
02What that means for how you measure
The person running that experiment drew the right conclusion and asked the right question: if AI recommendations are this volatile, what does ranking in AI search even mean, and is citation frequency the better metric to track instead of position?
Yes, and the reasoning is straightforward. A position from one run is a sample of size one drawn from a distribution that visibly moves. Frequency across many runs estimates the distribution itself, which is the thing that actually describes your visibility.
| Metric | Survives volatility? | Use it for |
|---|---|---|
| Position in a single answer | No | Nothing. Do not report this |
| Mention rate across a fixed prompt set, repeated | Yes | Trending your visibility over time |
| Share of voice versus named competitors | Yes | Competitive position, the useful one |
| Which sources were cited | Mostly | Deciding what to do next |
| A screenshot of one good answer | No | Reassuring an executive, briefly |
The practical rule: never make a decision from a single query result, and never let anyone else either. The screenshot showing your brand recommended first is real, and so is the one your competitor took twenty seconds later.
03Two gaps, and only one of them is actionable
Gap analysis in this context usually means one thing and should mean two.
The mention gap is the list of prompts where a competitor is named and you are not. It is what most tools report and it is diagnostic: it tells you a problem exists and roughly where.
The citation gap is which third-party sources the engine drew on to produce those answers, and whether you appear in them. This is the layer that tells you what to do, and it is the one an agency practitioner described as the actual roadmap.
“For each prompt I look at which sources the AI is citing in its answer. That's the actual roadmap to getting included. If the AI keeps citing a top painters in [town] listicle, getting onto that listicle matters more than another blog post.”
That example is from local services, and the structure transfers exactly to B2B software. If the engine keeps citing a category roundup, a review platform and two community threads when recommending tools like yours, then being absent from those four places is your visibility problem. Publishing a thirteenth blog post on your own domain does not address it.
The same practitioner describes step fourteen of their onboarding as finding where competitors are getting cited and they are not, and turning that into an outreach target list. That is the correct output of a gap analysis: a list of properties to earn a place in, not a list of pages to write.
04Where the prompt set comes from, and why it matters
Every number in a gap analysis is a function of the prompt list, which means the prompt list is the single largest source of bias in the whole exercise. An agency running this professionally is blunt about where theirs comes from.
“I build seed prompts from Google Search Console. This is the step people skip and it's the most valuable. GSC shows the real language people already use to find the business. No guessing at keywords, it's demand that's already there, just rephrased the way someone would ask an AI. ”
This solves a specific failure. A prompt set written from imagination encodes how your team describes the product, which is reliably different from how buyers describe the problem. That mismatch is where visibility gaps usually live, so a self-invented list tends to miss exactly the prompts you are losing.
The conversion is mechanical. Pull your queries, take the high-intent ones, and rewrite each as a full natural-language question. A two-word query becomes a sentence, because prompts are conversational and considerably longer than search queries.
Cover three shapes deliberately, since they behave differently: category prompts naming no brand, comparison prompts naming a competitor alongside you, and problem-first prompts that describe a situation rather than a product. Teams that only track category prompts miss the comparison losses that cost deals.
05Building a baseline that holds up
Given the volatility, a baseline needs both breadth and repetition. The numbers an agency uses in practice are a reasonable starting specification.
At least 50 prompts. Their reasoning: fewer than that and you are reacting to noise, while fifty or more gives enough coverage across services, locations, pricing and best-of style questions to see patterns and movement rather than one prompt bouncing around.
Several hundred individual chats behind the number. In one account they describe a benchmark drawn from roughly 700 chats across the prompt set, and explicitly note that the volume is what makes the starting visibility figure trustworthy rather than anecdotal.
Recorded per model, never averaged. Engines differ enough that a blended figure describes none of them. The same practitioner notes checking which model a client is already strongest in as the starting line, and found ChatGPT to be the strongest for one local client, which they flagged as less common than expected.
Frozen and versioned. Once the list exists, changing it resets your ability to trend. Add prompts as the account grows, but record when you did, because a jump in mention rate the same week you added fifteen easier prompts is not a result.
06Check you can be crawled before anything else
One step in that agency’s onboarding deserves to be first in yours, because it invalidates everything downstream if it fails.
They verify that every AI bot can actually reach the site, and report that bot-fighting settings on common security platforms frequently block AI crawlers. Their recommendation for a small business that is not a publisher is to turn that off, on the straightforward grounds that you want the AI to be able to read you.
This matters because a blocked crawler produces exactly the same symptom as bad content: absence from answers. Teams spend quarters on content when the problem was a checkbox. It also compounds with the rendering issue we documented in the server logs analysis, where several major AI crawlers found zero pages linked only by JavaScript.
The agency also reports AI crawler activity to clients as a leading indicator, showing bot visit growth over time. That is a reasonable early signal: crawler attention rises before citation does, so it gives you something honest to report while the visibility number is still flat.
07Reading the citation gap properly
Once you have the cited-source list, the analysis is a sorting exercise rather than a strategy exercise.
| Source type in the citation list | Realistic action | Effort |
|---|---|---|
| A category roundup or listicle | Get evaluated and included | Low to medium |
| A review platform profile | Claim and complete it, drive reviews | Low |
| A community thread | Participate genuinely, do not astroturf | Medium, ongoing |
| A competitor's own comparison page | Publish your counter-comparison | Medium |
| An encyclopedic reference | Long-term, often not realistic | High |
Two practical notes from practitioners. The first: a thin or incomplete profile on a major review platform is a citation that never happens, which makes completing those profiles one of the cheapest available wins. The second, from the agency AMA: they stopped looking at backlinks some time ago and spend their time on brand mentions instead, which is a meaningful reallocation of the same outreach effort.
On community sources, be careful. Communities are heavily cited, and they are also the fastest place to damage a brand by participating badly. Genuine participation over time works; posting to be cited does not, and is usually obvious to both readers and moderators.
One sequencing note that saves wasted effort. Work the sources that already appear in your citation list before chasing new ones. A roundup the engine already cites for your category is a known-good target; a publication you think ought to matter is a guess. The list the engine hands you is the only prioritisation input that comes with evidence attached, which is the whole reason to run the analysis rather than rely on judgement.
And re-run the citation pass quarterly rather than trusting last quarter’s list. Source weighting shifts fast enough that a target list built six months ago may point at properties the engine has since stopped leaning on, which we documented in the citation share benchmarks.
08Visibility on ChatGPT, Perplexity and AI Overviews is three different problems
A single AI visibility number is an average across engines that behave differently enough to need separate strategies. Published citation data makes the divergence concrete.
ChatGPT cites generously but links sparingly. Analysis of over a million citations found roughly 31% of URLs cited more than twice per answer, while a separate benchmark put ChatGPT’s rate of including external source links near 50% of responses. So it references sources heavily inside the answer while frequently naming brands with no clickable link, which means citation-based tracking systematically under-reports your ChatGPT presence.
Perplexity concentrates ruthlessly. The same dataset found 64% of URLs never cited at all, with the top 6% producing just under half of all citations, against a link rate around 96.5%. Perplexity is therefore the easiest engine to measure and the hardest to break into: either you are in the winning set or you are invisible.
Google AI Overviews are conservative and closest to organic. More than nine in ten URLs are cited less often than they are retrieved, and AI Overviews draw heavily on pages already ranking in traditional search. This is the engine where classic SEO investment transfers most directly, which is also why a gap here usually indicates an ordinary ranking problem rather than a new one.
| Engine | Behaviour | What a gap here means |
|---|---|---|
| ChatGPT | Cites generously, links about half the time | Track mentions, not just linked citations |
| Perplexity | Concentrates on a few sources, links almost always | You are outside the winning source set |
| Google AI Overviews | Conservative, overlaps heavily with organic | Usually a conventional ranking problem |
The practical consequence for gap analysis: prioritise by where your buyers actually are rather than trying to close every gap. Going deep on one or two engines and reaching a meaningful mention rate beats spreading across six and reaching a trivial one on all of them. The full engine-by-engine breakdown is in our ChatGPT tracker analysis and the pricing implications of tracking several engines at once are in cost per tracked prompt.
09Tracking sentiment in AI answers, carefully
Sentiment is the third thing teams want to track after presence and position, and it is the one to handle most cautiously.
The volatility finding applies here too, and arguably more strongly. If the recommended brand list changes between identical asks, the descriptive language around your brand will vary at least as much. A sentiment score derived from a handful of runs is measuring sampling noise.
What is worth tracking instead is specific claims. Not whether the tone was positive, but whether the engine repeatedly asserts something factual about you: a price point, a missing integration, a category placement. Recurring factual claims are stable enough to act on, and unlike a sentiment score they point at a fixable cause, usually an outdated third-party source.
Practitioners also report that consistency across your public profiles matters for how confidently engines describe you, with conflicting information across your site, review platforms and professional profiles making it harder for a model to determine which version is correct. We have not tested that claim directly, but keeping your own descriptions consistent is cheap and carries no downside.
The cheapest version of this check takes ten minutes: ask each engine to describe your company, and compare the answer against your own positioning. What you are looking for is not tone but specific errors, since a wrong price or a wrong category is both traceable to a source and worth fixing.
10The vanity metric problem
A marketer evaluating this whole category laid out the complaint that most gap-analysis reporting eventually runs into.
“Most GEO/AEO tools still seem focused on vanity metrics like you appeared in 12 prompts this week. As a marketer, I care less about mentions and more about whether those mentions create demand, clicks, branded searches, leads, and revenue.”
That is fair, and the honest position is that the connection to revenue is currently weak. Click-through from AI answers is low, attribution is poor because buyers frequently search your brand name separately afterwards, and no tool has solved this.
The workable compromise is to pair the visibility metric with one downstream proxy you already have: branded search volume. If AI visibility work is producing anything, the mechanism runs through people hearing about you and then looking you up. Branded impressions in Search Console are imperfect and first-party, which beats a vendor dashboard number with no outcome attached.
Report both, and be explicit that mention rate is a leading indicator rather than a result. That framing survives the meeting where somebody asks what it is worth, which a mention count on its own does not.
11The quarterly routine
Assembled from the practitioner processes above, sized for a team without a dedicated analyst.
Once, at the start. Verify AI crawler access. Build the prompt set from Search Console, at least 50 prompts across the three shapes. Name your real competitors rather than aspirational ones. Freeze and version the list.
Monthly. Run the set across the engines your buyers use, repeatedly enough to average out the volatility, recorded per model. Log mention rate and share of voice against the named competitor set. Log the cited sources.
Quarterly. Do the actual gap analysis. Which prompts do competitors win consistently across runs, not once. Which sources appear in those answers. Turn that into an outreach list rather than a content calendar.
Continuously. Work the source list. Review profiles, roundups, community presence, comparison pages. This is the part that moves the number, and it is the part that is not a dashboard.
One closing note from the agency AMA that is worth keeping, because it is the least glamorous sentence in this entire research set and probably the truest: this work is mostly disciplined fundamentals plus reading the citations carefully, not prompt tricks.
A final caution on expectations. Because the underlying answers move run to run, improvement shows up as a trend across months rather than as a step change you can point at in a week. Teams that check once a fortnight and conclude nothing worked are reading noise, in exactly the same way that teams celebrating a single good screenshot are. Set the review cadence to quarterly and hold the prompt set still in between, and the signal becomes legible.
Work the source list, not the dashboard
Linkeddit Compete tracks the community and review conversations that answer engines cite most heavily when recommending software, grades what changed, and returns a weekly brief. It is aimed at the citation gap rather than the mention count.
12Frequently asked questions
Frequently asked questions
How consistent are AI brand recommendations?+
Much less than most reporting assumes. One team ran the same buying-intent question repeatedly against the same model seconds apart with no prompt changes, and reported a different top brand 4 times out of 10, with the full recommended list overlapping only about half the time between asks. Even with randomness settings reduced as far as possible, the top pick still changed roughly one time in four. Markets with a clear dominant player were the exception.
What is a brand mention gap analysis in AI search?+
It is the comparison of which prompts your competitors appear in and you do not, run across a fixed prompt set and repeated over time. The useful version has two layers: the mention gap, meaning prompts where a rival is named and you are not, and the citation gap, meaning which third-party sources the engine drew on to produce those answers. The second layer is the actionable one because it tells you what to earn rather than what you lack.
How many prompts do you need for a reliable baseline?+
At least 50, and run repeatedly rather than once. An agency running this professionally described tracking a minimum of 50 prompts, on the grounds that fewer means reacting to noise, and building baselines from several hundred individual chats so the starting number is trustworthy rather than anecdotal. Given the volatility findings, repetition across runs matters as much as breadth across prompts.
Where should prompts come from?+
Your own Search Console data, not your imagination. An agency practitioner called this the step people skip and the most valuable one, because Search Console shows the real language people already use to find you, which you then rephrase as natural-language prompts. That is demand that already exists rather than guesses about how someone might ask, and it removes the biggest source of bias in a self-selected prompt set.
Should you track position or citation frequency in AI answers?+
Frequency, given the volatility. If the top recommendation changes on repeated asks of an identical question, a position number from a single run describes that run and nothing more. Mention frequency across many repetitions of a fixed prompt set is stable enough to trend, which is why practitioners increasingly report share of voice and mention rate rather than rank.
Can AI crawlers be blocked without you realising?+
Yes, and it is a common finding. An agency running these audits reported that bot-fighting settings on common CDN and security platforms frequently block AI crawlers, which makes verifying crawler access an explicit step in their onboarding. A brand that cannot be crawled cannot be cited, so this belongs at the start of a gap analysis rather than after months of content work.