AI Visibility · Measurement
AI Recommendation Monitoring: Mentioned vs Chosen
Search this topic and you get social listening tools that count how often your name appears on the web. That is a different question from whether an AI assistant recommends your product when a buyer asks which one to use, and the second question is the one with revenue attached.
Key takeaways
- Being mentioned, being cited and being recommended are three different outcomes. Most tools detect the first because string matching is easy.
- An absolute visibility score can rise for every tracked brand at once. Rank against a named competitor set is the version that carries information.
- Position inside the answer matters and varies by engine, with some surfacing brands near position one or two and others pushing them deeper.
- Recommendations are volatile enough that single-run measurement is noise. One team saw the top brand change 4 times in 10 on an identical repeated question.
- Demand three capabilities in a demo: competitive rank, position within the answer, and the cited sources behind it. Many tools cannot show all three.
01Mentioned, cited, recommended: three different things
The category calls all of this brand monitoring, which hides a distinction that decides what you should do next.
| Outcome | What happened | What it tells you |
|---|---|---|
| Mentioned | Your name appeared somewhere in the answer | The model knows you exist in this category |
| Cited | A link to your domain supported the answer | Your content was usable evidence |
| Recommended | The model put you forward as the answer | You won the query |
The gap between rows one and three is where deals sit. A model can mention your product in a sentence explaining why a buyer might prefer a competitor. String matching records that as a mention, and your dashboard shows a green number for an answer that actively cost you the deal.
That is not a hypothetical failure. It follows directly from how these tools work: they run a prompt, parse the response for brand strings, and count. Unless a tool evaluates the role your brand played in the sentence, appearing and winning are indistinguishable to it.
The practical consequence is that a rising mention count is ambiguous evidence. It is consistent with winning more often and equally consistent with being named more often as the option someone rejected. Without the role your brand played in the sentence, the direction of the change is unknowable, which is an uncomfortable property for a metric people report monthly.
02Why most tools measure the easy one
There are two structural reasons the category defaults to mention counting, and neither is laziness.
Mentions are cheap and unambiguous to detect. Searching a response for your brand name is deterministic. Judging whether the model recommended you requires interpreting the answer, which is a harder, slower and less reliable operation.
Citations are only visible about half the time. A benchmark published in February 2026 put ChatGPT’s rate of including external source links near 50% of responses, against roughly 96.5% for Perplexity. So a tool tracking linked citations is blind to about half of ChatGPT’s output, which pushes vendors back toward mention counting as the only signal available across engines.
The result is a category-wide measurement bias toward the outcome that is easiest to detect rather than the one that matters most. We covered the structural limits this creates in the ChatGPT tracker analysis.
03Rank, not score
The sharpest correction to how this gets measured came from a practitioner who had tested several tools and distilled what separates a useful one.
“Visibility rank not just visibility score: score can go up for everyone at once. Rank tells you where you actually stand relative to competitors. Big difference.”
That is the whole argument in two sentences. An absolute visibility score has no denominator. If engines start citing more sources generally, or if your category simply gets discussed more, every tracked brand’s number rises while competitive position is unchanged.
Rank against a named competitor set does not have that problem, because it is relative by construction. It also matches the question executives actually ask, which is never what is our visibility score but always whether we are ahead of the two companies we lose to.
04Position inside the answer, which almost nothing reports
A second distinction sits underneath recommendation: where in the answer you appear.
Being named first in a two-option answer and being named eighth in a list of ten are recorded identically by a mention counter, and they are not remotely the same commercial outcome. Eighth in a list is closer to absent than to recommended, because a reader scanning an answer rarely evaluates the full set.
Engines differ here in ways worth knowing. Spotlight’s benchmark reported that Perplexity, ChatGPT and Grok tend to surface brands near position one or two within an answer, while Claude pushes them deeper. If your buyers concentrate in an engine that surfaces few brands prominently, position is close to binary and being outside the top two is effectively invisibility.
The practical ask in a demo is simple: show me, for one of my prompts, not just whether we appeared but where, and who appeared above us. A tool that cannot answer that is reporting presence, not competition.
05The measurement problem underneath all of this
Before trusting any recommendation metric, one finding has to be accounted for, and it undermines single-run reporting entirely.
“It named a different top brand 4 out of 10 times. The full list of recommended brands only overlapped about half the time from one ask to the next.”
That team asked an identical buying question repeatedly, seconds apart, with no prompt changes. Their exception is instructive: markets with a clear dominant player were stable, and contested categories were close to a coin toss.
Two consequences for recommendation monitoring specifically. Recommendation is the most volatile of the three outcomes, because it is a single winner rather than a set, so it needs the most repetition to measure. And a vendor demonstrating your strong recommendation position in a live demo is showing you one draw from a distribution, which is worth roughly nothing.
The workable approach is frequency across many repetitions of a frozen prompt set, reported per engine, which we set out in the gap analysis method.
06The three capabilities to demand
Rather than a tool ranking, which ages badly in this category, evaluate on three capabilities. Ask for all three on one of your own prompts, live.
| Capability | The question it answers | Why vendors struggle |
|---|---|---|
| Competitive rank | Are we ahead of the two companies we lose to? | Requires a named competitor set, not just your brand |
| Position in answer | Were we the recommendation or the eighth option? | Requires parsing structure, not string matching |
| Cited sources | What evidence produced this answer? | Only available on roughly half of ChatGPT responses |
The third is the one that converts monitoring into action. Knowing you were not recommended tells you there is a problem. Knowing the answer drew on three community threads and a comparison page that all favour a competitor tells you what to do about it.
Two further questions worth asking, because they materially change reported numbers. What exactly counts as a mention, meaning whether product names, aliases and bare domain references are resolved. And how many repetitions sit behind a reported figure, given the volatility above. Vendors differ enormously on both and rarely volunteer either.
08Diagnosing which problem you actually have
Before buying anything, work out which of four situations you are in. They look similar in a dashboard and need completely different responses.
| Situation | Symptom | Where the problem is |
|---|---|---|
| Not mentioned at all | Absent from category prompts | The model does not associate you with the category |
| Mentioned, never recommended | Present but always as an also-ran | Positioning, or thin third-party evidence |
| Recommended in some engines only | Strong in one, absent in another | Source mix differs by engine |
| Recommended but not cited | Named with no link to you | Fine, and largely unmeasurable |
The second row is the most common and the most misdiagnosed. Teams see the mention count and conclude visibility work is succeeding, while every one of those mentions is a sentence positioning them as the cheaper alternative to someone else.
The fourth row deserves a word because it causes unnecessary alarm. Being recommended without a link is a good outcome that your tooling will under-report, given that roughly half of ChatGPT responses carry no external link. Do not optimise away from a result that is working because a dashboard cannot see it.
A closing word on why this distinction is worth the trouble. Every metric in AI search is currently weak, and the temptation is to pick whichever one moves and report that. Mention counts move readily, which is precisely why they get chosen. Rank against a named competitor set moves less and means more, and a team reporting the harder number keeps its credibility when somebody eventually asks what any of it bought.
09Different problems, different fixes
The reason to separate these outcomes is that the remedies do not overlap much.
Not mentioned at all is usually an association problem. The model does not connect you to the category, which is fixed by presence in the places that define the category: roundups, review platforms, community discussion. This is off-site work, not on-site work.
Mentioned but never recommended is usually an evidence problem. Something in the source material positions you as the secondary option. The fix is finding which sources produce that framing and addressing them, which requires the cited-sources capability above.
Engine-specific gaps are source-mix problems. Engines lean on different source types, so a gap in one usually means you are absent from the source category that engine favours rather than absent generally. The engine-by-engine behaviour is in the citation benchmarks.
Notice that two of the three fixes are off-site. That is consistent with the wider finding that AI citation concentrates heavily in a small set of third-party domains, and it means recommendation monitoring is mostly useful as a way of prioritising off-site work rather than on-site publishing.
10What actually moves a recommendation
Monitoring is only worth running if something can change as a result. Four levers show up consistently, ordered here by how much evidence supports them rather than by how satisfying they are to work on.
Presence in the sources the engine already cites. The strongest lever and the least glamorous. Research into AI citation puts the top fifteen domains at roughly 68% of consolidated citation share, against about 20% for the top fifteen in Google organic. Most of those domains are not yours, which makes earning a place in them the highest-leverage available action.
Comparison content that names the competitor. Buying prompts are frequently comparative, and comparison pages are heavily used when engines answer them. A competitor with a page comparing themselves to you, unanswered, is supplying the framing the model repeats.
Extractability of your own strongest pages. A page that ranks but buries its answer in a paragraph of preamble is harder to quote than one that states the answer directly under a descriptive heading. This is cheap to fix and only helps where you are already being retrieved.
Consistency across your public profiles. Practitioners report that conflicting descriptions of a company across its own site, review profiles and professional pages make it harder for engines to describe the brand confidently. We have not tested this directly, and keeping descriptions consistent costs almost nothing regardless.
Notice what is absent: nothing on this list is a prompt trick, and three of the four are off-site or editorial rather than technical. That matches every other line of evidence we have gathered on AI visibility, and it is the main reason recommendation monitoring is worth running at all, since it tells you which of these four to spend on first.
11Reporting this without overclaiming
The hardest part of this work is not measurement, it is presenting it to someone who will reasonably ask what it is worth.
Be direct about the limitation. Click-through from AI answers is low, and buyers frequently search your brand name separately afterwards, which analytics records as direct or branded traffic with no attribution path back. You can be winning recommendations and be unable to prove the revenue effect.
Three practices make the report survive scrutiny anyway.
Lead with rank, not score. Are we ahead of the two companies we lose to, across a frozen prompt set, per engine. That framing survives the question a score does not, which is what happened when everyone else improved too.
Pair it with branded search volume. If recommendation work is producing anything, the mechanism runs through people hearing about you and then looking you up. Branded impressions are first-party, imperfect, and a better companion metric than any vendor number with no outcome attached.
State the volatility in the report itself. Including the variance alongside the figure protects you twice: it prevents a good month being read as a permanent gain, and it prevents a bad month being treated as a failure of the programme. A number presented without its noise band invites exactly the wrong conversation in both directions.
12Measuring it yourself, which is more feasible here than elsewhere
Recommendation monitoring is unusually amenable to manual measurement, because the thing you are recording is simple even though it is volatile.
Write ten buying prompts. The questions a buyer asks when choosing, not describing. Best tool for a use case, alternatives to a named competitor, which of these should we pick.
Run each five times, per engine. Five is the minimum that makes the volatility visible. This is tedious and it is the entire method.
Record three columns, not one. Were you named, in what position, and who was named first. The third column is your competitive rank and it is the number worth trending.
Record it in a spreadsheet rather than a document. The value of this exercise is comparability across quarters, and prose notes stop being comparable almost immediately.
Tally the cited sources separately. Across all fifty runs, which domains keep appearing. That list is the actual output of the exercise and it does not require any tool.
An afternoon produces a defensible baseline and, more usefully, tells you whether your category is stable or volatile, which determines how much repetition any future measurement needs. Teams that skip this and buy a tool first end up with a dashboard they cannot interpret, because they have no sense of how much of the movement is real.
Repeat quarterly rather than continuously. Given the volatility, monthly manual runs mostly measure noise, while a quarterly comparison against a frozen prompt set is long enough for a real change to separate from the variance, and short enough that a competitor making a deliberate move shows up before it compounds.
The sources behind the recommendation
Linkeddit Compete tracks the community and review conversations answer engines lean on when recommending software, and grades what changed week to week. It is aimed at the evidence layer that decides who gets recommended, rather than at counting mentions.
13Frequently asked questions
Frequently asked questions
What is the difference between AI brand mention monitoring and recommendation monitoring?+
A mention means your name appeared somewhere in the answer, possibly in a list of eight options or in a sentence explaining why someone chose a competitor. A recommendation means the model put you forward as the answer. They are different outcomes with different fixes, and most tools in this category measure only the first because it is far easier to detect with string matching.
Why is a visibility score not enough?+
Because an absolute score has no denominator. As one practitioner put it, a score can go up for everyone at once, whereas rank tells you where you actually stand relative to competitors. If engines begin citing more sources generally, every tracked brand’s score improves while nobody’s competitive position changes, which makes the number look like progress when nothing has moved.
Does position within an AI answer matter?+
It appears to, and it varies by engine. Spotlight’s February 2026 benchmark reported that Perplexity, ChatGPT and Grok tend to surface brands near position one or two in an answer while Claude pushes them deeper. Being named eighth in a list is closer to being absent than to being recommended, so a tool that reports presence without position is compressing away the part that matters.
Can you reliably measure AI recommendations at all?+
Only across repeated runs. A team that asked an identical buying question repeatedly, seconds apart, reported a different top brand 4 times out of 10 and roughly 50% overlap in the recommended list between asks. In contested categories the top recommendation was close to a coin toss, which means any single-run recommendation report is measuring noise rather than position.
Which tools actually track recommendations rather than mentions?+
Look for three specific capabilities rather than brand names, because the category changes fast: competitive rank against a named competitor set rather than an absolute score, position within the answer rather than binary presence, and the cited sources behind each answer. Ask a vendor to show all three on one of your real prompts during a demo. Many cannot.
Why do social listening tools rank first for these searches?+
Because they have far larger domains and have owned brand monitoring terminology for a decade. The mismatch is real: social listening measures conversation volume about your brand across the web, which is a different job from measuring whether a model recommends your product when a buyer asks which tool to use. Both are called brand monitoring and they answer different questions.
07Why social listening tools dominate these searches
Search for brand monitoring software in this context and the first results are social listening platforms with large domains and a decade of ownership over the terminology. That is a vocabulary collision rather than a recommendation.
Both are legitimate and they answer different questions. The mistake is buying one expecting the other, which happens frequently because both are sold as brand monitoring. If your question is whether an assistant recommends your product when a buyer asks, a social listening platform will not answer it regardless of how good it is at its own job. We covered what those platforms do well, and where practitioners find them weak, in the social listening guide.
One practical note if you already own a social listening subscription: point it at competitor names rather than cancelling it. Conversation volume about your competitors, and the complaints inside it, is useful input to the evidence problem described below, even though it will never tell you what a model recommends.