Answer Engine Optimization
Why Every AI Engine Names a Different Brand
A practitioner ran the same brand queries through four engines across ten verticals and got near-zero overlap, plus engines citing a brand's own page as the source for recommending it. Both findings are real. Only one is about the engines.
Key takeaways
- Two different things produce cross-engine disagreement: engines read different source pools, and each engine also disagrees with itself between runs. A four-tab comparison cannot tell you which one you are looking at.
- A recommendation whose only citation is the recommended vendor's own page is circular. The citation restates the claim rather than supporting it, and that is a fact about the category's evidence layer, not about the vendor.
- Agreement between engines is not validation. Check whether the agreement came from the same one or two documents before you read it as consensus.
- For a buyer, the useful output is the union of the four answers as a candidate list, then the citations as the filter. For a brand, the useful metric is appearance rate per engine over repeated runs, never rank.
01Why do four AI engines recommend four different brands?
Because they are not answering from the same evidence, and because none of them answers the same way twice. Those are two separate causes, they stack, and almost every guide on this topic explains the first while ignoring the second, which is how four screenshots get mistaken for a finding.
The retrieval half is well documented. Yext analysed 17.2 million citations across the four engines and reported that Claude cited user-generated content at two to four times the rate of the others, while verified structured listings made up 54.53% of distinct citation sources overall. A separate citation audit found that only around 11% of domains cited by ChatGPT are also cited by Perplexity. Different pools, different weightings, different answers. That much the ranking pages all say, and it is true.
“This is why traditional SEO metrics alone won’t explain LLM visibility anymore. Different models clearly trust different authority signals.”
What that framing leaves out is the part that changes how you read your own test. Each engine is a sampler. It does not hold a ranked list of vendors and read it out. It generates a plausible list, and the list is different next time. So when four engines give you four answers, some unknown share of that gap is the engines disagreeing with each other and the rest is each engine disagreeing with itself. Until you separate those, you have not measured anything about the engines. You have taken four samples of size one.
02How much of the disagreement is real, and how much is noise?
More of it is noise than anyone expects. The largest public study on this asked 600 volunteers to run twelve brand-recommendation prompts through ChatGPT, Claude and Google’s AI a combined 2,961 times, then measured how often the same list came back. The finding was that there is less than a 1 in 100 chance of seeing the same list of brands twice, and closer to 1 in 1,000 of seeing the same list in the same order.
Sit with what that does to the comparison you ran. If a single engine will not repeat its own list once in a hundred tries, two engines producing different lists is the expected outcome even if their view of your category were identical. The cross-engine gap you measured contains the within-engine gap you did not. Its author, a self-described sceptic of the category, still concluded that appearance frequency across many runs is trackable while ranking position inside an answer is not.
Practitioners arrived at the same place from the other direction. The metric that keeps getting reinvented in these threads is persistence: not where you appeared once, but whether you appear again tomorrow.
“That is how frequently is a brand cited for a specific search phrase, overtime. Or in other words, if the same search phrase is made every day for 10 days, does the brand in question get cited every day.”
That is the right unit. Stability first, then comparison. An engine that names you in eight runs out of ten and an engine that names you in two is a real difference between engines. An engine that named a competitor at 9am and you at 2pm is one engine, sampling.
03Why does Perplexity cite the recommended brand's own website?
Because a retrieval-based answer has to attach a document to the claim, and in a category with no trusted independent comparison, the closest-matching document is frequently the vendor’s own page. The engine is not endorsing the source. It is filling a required slot with the best available match, which was written by the company being recommended.
The result is a circular answer. The citation does not corroborate the recommendation, it restates it in the vendor’s own words. A buyer skimming the answer sees a link and reads it as evidence. There is no evidence there. Read it as a category diagnostic rather than a scandal; the fuller version of that argument is in branded versus discovery prompts.
Two related patterns look identical from the outside and need different responses. The first is substitution: the engine reaches for the biggest aggregator it can find instead of any vendor at all.
“I can talk from a user perspective: I always get recommendations of bigger brands on ChatGPT. When i check the links, it's using Amazon as a reference.”
The second is mundane and probably more common. One engine cannot read you at all because its crawler is blocked, so it falls back to whatever it can reach, and you record a cross-engine insight that is actually a robots file.
“they have a tool that will check if your site can be reached by the AI indexing bots, as they're so often accidentally blocked.”
Check the boring cause first. It costs ten minutes and it invalidates a week of analysis if you skip it.
04What should you do when you are the buyer and the engines disagree?
Almost everything written about this question is written for the brand trying to get recommended, and very little for the person on the other side, who asked an assistant which tool to buy and got four incompatible answers. That reader has a better protocol available than picking the engine they like most.
Start by refusing the shortlist. Four engines sampled four times from overlapping pools of candidates, so the useful object is the union of those answers, not the intersection and certainly not the top item in any one of them. Then use the citations as the filter, because that is the only part of an AI answer you can check.
| What you see | What it usually means | What to do |
|---|---|---|
| A recommendation citing only the vendor's own page | Nobody independent wrote the comparison, so retrieval used the vendor's marketing | Treat it as unverified until you find a third-party account |
| The same brands everywhere, different sources | They are prominent enough to sit in every model's memory | Prominence, not fit. Check whether any serve a company your size |
| The same brands and the same one listicle | A single document is deciding the category | Read it, check who wrote it and whether it is affiliate-funded |
| Four completely different answers | Thin independent evidence, or a category too new to have settled | Expect to do the comparison yourself |
| A confident claim with no citation at all | The model is answering from memory, which can be stale or wrong | Verify before repeating it |
One more move, and it is free: ask again. The consistency research puts it bluntly, that if you do not like an answer you can run the prompt a few more times and get a different one. That cuts both ways. No single answer deserves the weight buyers give it, and three runs tell you far more than one about what the engine associates with your problem.
05What does the disagreement tell a brand about its own category?
It tells you where the independent evidence layer is thin, which is more actionable than any per-engine tactic. Categories with real third-party infrastructure, review platforms people use and comparisons written by someone with nothing to sell, produce far more agreement between engines. Categories without it scatter, and self-citation fills the gap.
That reframes the finding. Wild disagreement across four engines is not evidence that the engines are broken. It is evidence that nobody credible has written the comparison your buyers need, which is a gap you can fill and a competitor cannot block. The practitioners tracking this land on the same two levers.
“Content specificity and novelty, followed by brand sentiment and reputation on non-owned platforms, are the two factors that move the needle in my experience and testing.”
Note the direction of both. Neither is a per-engine trick. Specific, verifiable claims and evidence that lives somewhere you do not own are what every retrieval path can find, which is why they narrow the gap between engines instead of winning one of them. The corollary is uncomfortable for content teams: publishing more on your own domain feeds the engines that already lean on owned pages and does nothing for the one reading forums and review sites.
Agreement deserves suspicion for the same reason. If all four engines name the same three vendors and cite the same comparison article, your category is being decided by a single document, which is fragile for whoever is winning and tractable for everyone else. For a measured version of how little the engines overlap, see our run of 21 buying questions across five engines on one day.
06How do you measure cross-engine disagreement properly?
Fix the prompts, repeat them, report rates with denominators. In that order, because each step is worthless without the one before it.
- Fix a prompt set and freeze it. Fifteen to twenty five questions phrased the way a buyer describes a problem, not the way you describe your product. Rewording between runs makes every comparison meaningless.
- Repeat within each engine before comparing across. Three to five runs per prompt per engine is the floor. Below that you cannot tell an engine that dislikes you from an engine that sampled around you.
- Report appearance rate, never rank. Named in 7 of 10 runs survives scrutiny. Ranked second is not a number at all.
- Log the cited source every time. Whether you appeared says there is a problem. Which source got cited says where the fix lives.
Two numbers come out of that. Within-engine stability tells you whether the engine has a settled view of your category. Cross-engine agreement, computed only on prompts that are stable within each engine, tells you whether the engines are reading the same evidence. Reporting the second without the first is the most common mistake in this discipline, and it is why two audits of the same brand in the same week can disagree completely.
Hold the vendor selling you the measurement to it too.
“so if a tool can’t clearly explain where its data comes from or how it’s measured, that’s usually a sign to be careful.”
07What this does not tell you
Three honest limits, because this category is full of confident claims that do not survive contact with a denominator.
Nobody measures real buyer sessions. Every tracker, ours included, runs its own prompts and infers behaviour from them. Real people write prompts that barely resemble each other or the ones you tested, which the same research showed when 142 volunteers describing one intent produced almost no textual overlap. Appearance rate on your prompt set is a proxy and should be labelled as one.
The specific agreement numbers circulating are not verified. One practitioner in these threads reported that fifty prompts across three engines agreed on only 21% of recommended brands. We could not reproduce it, no methodology was published with it, and nobody in the thread replicated it. Treat it as an anecdote with a number attached.
Some of this may not be worth your attention yet. The strongest counter-argument in these threads is not that the measurement is wrong, it is that the surface is small.
“SEO is AEO. While LLMs are suggesting more, they don’t (yet) represent a significant portion of search. Maybe 5-10%.”
That is a fair position and we will not pretend otherwise. The narrower case for doing the work now is that the answer to a problem description is a shortlist you never see, and unlike a search ranking, nobody tells you that you were left off it. If the engines in your category disagree wildly today, that is the window: the evidence layer they will settle on has not been written yet.
Measure the disagreement instead of screenshotting it
Frequently asked questions
Why do ChatGPT, Claude, Perplexity and Gemini recommend different brands for the same query?+
Two causes stack. The first is retrieval: each engine reads a different pool of documents and weights different evidence, so they are not disagreeing about the same facts, they are reading different sources. The second is sampling: each engine also disagrees with itself, returning a different list almost every time the same prompt runs. One side-by-side comparison cannot separate the two, which is why a screenshot of four tabs is an anecdote rather than a measurement.
Why does Perplexity cite the recommended brand's own website?+
Because a retrieval-based answer needs a document that states the claim, and in categories with no trusted independent comparison, the best-matching document is often the vendor's own page. The citation then restates the recommendation instead of corroborating it. Read it as a signal that your category has a thin third-party evidence layer.
Does it mean anything when all four engines agree?+
Less than it looks. Agreement on brands with disagreement on sources usually means those brands are prominent enough to sit in every model's memory. Agreement on brands and on sources means one or two documents are deciding your category. Check which sources produced the agreement before reading it as validation.
How many times should you run a prompt before trusting the answer?+
More than once, and the honest number is uncomfortably high. The largest public consistency study ran twelve prompts sixty to a hundred times each and concluded that only frequency of appearance across many runs is stable enough to track, while ranking position is noise. If you cannot run that volume, run each prompt three to five times and report a rate rather than a position.
As a buyer, which engine should you trust when they disagree?+
None of them as a shortlist. Treat the union of the answers as a candidate list, then check the citations. A recommendation supported only by the recommended vendor's own page has no independent evidence behind it. One supported by a review platform, a practitioner write-up or a comparison you can read yourself is worth an hour.
Is cross-engine disagreement a problem you can fix?+
You can narrow it, not close it. Consistent facts about your product across your own site, directories, review platforms and third-party comparisons give every retrieval path the same story to find. The engines will still sample differently, so the realistic target is a higher appearance rate in each engine, not an identical answer from all four.