Answer Engine Optimization · Research
AI Citation Sources: Engines Barely Cite the Same Sites
Most AEO advice treats AI search as one channel. Practitioner studies put the domain overlap between ChatGPT and Perplexity near 11 percent. This covers how to find the sources that decide your category, and why the numbers in every published study, including the ones here, expire fast.
Key takeaways
- A practitioner who ran 100,000 prompts through ChatGPT and Perplexity reported roughly 11 percent domain overlap. Optimising for one engine says little about the others.
- The engines differ because their retrieval differs: ChatGPT draws on a search index, Perplexity crawls continuously, Claude leans on training data, and Google AI Overviews still tracks organic results closely.
- Reported figures for the same fact disagree wildly. Perplexity's Reddit share appears as 46.7 percent in one analysis and up to 20 percent in another. Treat all of them as directional.
- Citation patterns are unstable. Reddit's share of ChatGPT citations reportedly fell from around 3.8 percent to near 0.5 percent in about two weeks in August 2026.
- The durable practice is not memorising a source list. It is measuring your own category on a schedule, because the list changes without notice.
01The engines do not share a source list
The working assumption behind most AEO advice is that AI search is one channel. The citation data says it is several.
A practitioner working on this professionally described running a large comparison and finding far less common ground than expected:
“At work, we ran 100,000 distinct prompts through ChatGPT and Perplexity and found only 11% of cited domains overlap. That means 89% of citations come from completely different sources depending on which engine someone uses.”
The same 11 percent figure circulates in AEO communities attributed to a 2026 analysis, which is either corroboration or the same number travelling. We could not establish which, and are reporting it as a practitioner claim rather than an established finding.
What is corroborated across several independent accounts is the direction and the rough shape of each engine’s preferences:
| Engine | Reported source leanings | Confidence |
|---|---|---|
| ChatGPT | Encyclopedic and consensus sources, with Wikipedia repeatedly named. Business and tech press. | Consistent across accounts. |
| Perplexity | Community and discussion sources, Reddit most often named. Also YouTube and LinkedIn. | Consistent direction, contested magnitude. |
| Google AI Overviews | Closest to organic results. One report puts overlap with the organic top ten near 40 percent. | Single source, plausible. |
| Claude | Depth and structured content. One account reports it is more likely to cite bullet-pointed pages. | Single source, unverified. |
| Gemini | Video-heavy in some categories, with YouTube dominant in one client graph. | Single case study. |
Note the confidence column. Most of what circulates about per-engine citation preferences rests on one vendor study each, and vendors in this category have an obvious interest in the answer.
02Why retrieval architecture drives the difference
The divergence is not arbitrary. It follows from how each engine actually gets information at answer time. A practitioner guide summarising this put the mechanism plainly: the engines are not reading the same internet, because they have different retrieval architectures.
- An engine backed by a search index inherits that index’s authority signals, which is why encyclopedic and established-publisher sources appear so heavily.
- An engine that crawls continuously can surface recent discussion, which is why community sources rank higher in its citation mix.
- An engine leaning on training data has a knowledge cutoff and cites less live material, favouring what was well-represented in training.
- An engine built on top of organic search stays closest to conventional rankings, which is why traditional SEO carries over there better than anywhere else.
This is why the volume of citations per answer also differs so much. One analysis put Perplexity near 21 sources per response against roughly 7 for ChatGPT. That has a direct strategic consequence:
An engine citing three times as many sources is three times easier to appear in and considerably harder to dominate, because each citation carries less weight in the final answer. Entering the citation set is not the same as being the recommendation, which is the distinction covered in measuring AI search visibility.
03How to find your category's real sources
This takes about an hour and needs no tooling. The output is a ranked list of the domains that actually decide recommendations in your category, which is almost always different from the list you would have guessed.
A practitioner working on citation tracking described the manual method clearly, and the critical instruction is in the third sentence:
“Prompt with the buying question a customer would ask, like 'best tools for X', rather than your own brand name. You want the sources the model trusts when it decides who to recommend. After ten or twenty questions the same handful of domains keep appearing.”
The procedure:
- Write 15 to 20 unbranded buying questions. Category questions, not brand lookups. A brand-name query returns your own site and teaches you nothing.
- Run them in each engine that matters to your buyers. Perplexity exposes numbered citations directly. For ChatGPT, phrase questions so they require current information and then ask which sources were used.
- Log the domain behind every citation. Not the page. The domain is the unit that repeats.
- Count and rank. After twenty questions the distribution is usually stark: a handful of domains carry most of the citations, with a long tail appearing once each.
- Separate the reachable from the unreachable. Some of your top sources will be places you can realistically appear (review sites, community threads, industry write-ups). Some will not be. Only the first group is a work list.
04Why every published list expires
The most important thing to understand about citation source data, including everything on this page, is that it has a short shelf life.
The clearest recent example: practitioners tracking citation share reported that Reddit held roughly 3.8 percent of ChatGPT Search citations from mid-July into early August 2026, then fell to about 0.5 percent by 14 August. That is an approximately 86 percent relative decline within about two weeks, affecting a source that had been among the engine’s most-cited for two years.
One operator running per-brand citation tracking reported the same pattern in a live account, with Reddit citations at zero for a client in the current window and the engine leaning on primary and manufacturer sources instead. They were careful to caveat it as a single-account snapshot, which is the right level of confidence.
Two conclusions follow, and the second is the useful one. First, treat every most-cited-domains listicle as a dated snapshot rather than a structural fact. Second, the durable practice is not knowing the list, it is running the measurement. A team measuring its own category monthly would have detected the Reddit change within weeks. A team working from a published list would still be optimising for it.
05Building on the overlap set
Building separately for every engine is not affordable for most teams. The efficient structure is a foundation plus a layer.
The practitioner who ran the 100,000-prompt comparison drew the same operational conclusion, recommending starting with the sources that appear across engines and building platform-specific tactics on top of that base. Domains repeatedly named as appearing across multiple engines include Reddit, Wikipedia, YouTube, G2, and major business press.
| Layer | What it buys | Effort profile |
|---|---|---|
| Overlap sources | Broadest cross-engine carry from one piece of work. | Do first. Highest return per hour. |
| Your single most important engine | Depth where your buyers actually are. | Do second, chosen from your own measurement. |
| Remaining engines | Marginal coverage. | Only with budget left over. |
Choosing the second layer requires knowing which engine your buyers use, which self-reported attribution answers better than any benchmark. That method is covered in AI search ROI and attribution.
One caution on the overlap set: it is exactly as unstable as anything else here. The August ChatGPT change removed a major overlap source from one engine without warning. A foundation built on cross-engine sources is more robust than a single-engine bet, not permanently safe.
06Reading a competitor's citation sources
The same method points at a competitor and answers a question no rank tracker ever could: not that they beat you, but what the engine read before deciding they should.
Run the unbranded buying questions, and when a competitor is recommended, log the citations attached to that answer. Over twenty questions this produces their effective source set, and comparing it against yours produces a specific, checkable work list:
| What the comparison shows | What it means | Action |
|---|---|---|
| They appear in sources you are absent from. | A reachable coverage gap. | Get represented on those specific surfaces. |
| You appear in the same sources but are not recommended. | A framing problem, not a coverage problem. | The source describes them more usefully. Fix what it says. |
| They are cited from their own domain, you are not. | Their content answers the question directly. | Build the page that answers it, not a product page. |
| Neither of you is cited; a third party decides. | The category is mediated by a review site or community. | That surface is the battleground, not your website. |
The last row is the one that most often surprises teams, and it is where the work sits in categories dominated by review sites and community discussion. We cover reading those surfaces in competitor review analysis and the underlying question of why an engine prefers a rival in why ChatGPT recommends your competitors.
07The Reddit question, specifically
Reddit deserves its own section because it is simultaneously the most-cited community source across engines and the clearest recent example of how fast this can change.
The case for it: Reddit appears in nearly every cross-engine overlap list, and Perplexity in particular has been reported to lean on it harder than any engine leans on any single domain. One tracking programme covering more than 200 million prompts across five months described Perplexity citing Reddit in up to 20 percent of responses, against a peak near 5 percent for any single domain in ChatGPT.
The case against relying on it: the August 2026 drop in ChatGPT removed most of that engine’s Reddit citations in about a fortnight. If your entire AEO strategy had been Reddit presence, one engine went dark on you with no notice and no announcement.
| Reading | Wrong conclusion | Better conclusion |
|---|---|---|
| Perplexity cites Reddit heavily. | Post on Reddit and you will be cited. | Reddit threads in your category are a surface worth being discussed on, honestly. |
| ChatGPT stopped citing Reddit. | Reddit no longer matters. | One engine changed. Check what replaced it in your own measurement. |
| Reddit is in every overlap list. | It is a permanent foundation. | It was, in the window those studies covered. |
The honest position: community discussion is a genuinely important citation surface in most B2B categories, it is not controllable the way owned content is, and its weight varies by engine and over time. Being discussed accurately in the threads where your category is debated is worth real effort. Treating any single platform as a durable strategy is what the August data argues against.
We should disclose our own position here: Linkeddit operates a Reddit data pipeline, so we have an obvious interest in Reddit mattering. That is exactly why this section reports the decline as prominently as the case in favour.
08What to actually do about it
The findings above collapse into a short list of things worth doing, in order.
- Measure your own category before reading another study. Twenty unbranded questions per engine. Your list will differ from every published list, because those aggregate across categories that behave nothing like yours.
- Fix the reachable sources first. A review-site profile that is thin or outdated is usually the fastest fix available, and review sites appear in most categories’ top citation sets.
- Answer the question on your own domain. Engines cite pages that answer the buying question directly. Product pages rarely do; comparison and method pages often do.
- Re-measure monthly with the same questions. This is the only defence against the volatility described above, and it is cheap once the question set exists.
- Do not chase a single engine’s quirk. Any tactic aimed at one engine’s current preference is a bet that the preference persists. August 2026 is the argument against that bet.
- Keep a dated log of what you found. The value compounds only if you can compare this month against three months ago. A spreadsheet with a date column beats a dashboard you cannot export, because the whole point is the trend rather than any single reading.
When the re-measurement is the bottleneck
The twenty-question audit above is genuinely doable by hand and you should run it that way first. Answer Radar exists for step four: running the same frozen question set on a schedule across engines and recording which domains were cited each time, so a shift like the August Reddit change shows up in your own data rather than in somebody’s blog post two months later. We build in this category, so treat this as the disclosure it is.
09Frequently asked questions
Frequently asked questions
Do different AI engines cite the same websites?+
Much less than most people assume. A practitioner who ran 100,000 distinct prompts through ChatGPT and Perplexity reported only about 11 percent of cited domains appearing in both, a figure separately attributed to a 2026 analysis circulating in AEO communities. Treat the exact number as a practitioner report rather than a settled fact, but the direction is corroborated repeatedly: the engines run substantially separate source ecosystems, so being cited well in one says little about the others.
How do you find which websites ChatGPT cites for a keyword?+
Ask the buying question rather than your brand name, then read the sources. Perplexity is easiest because every answer carries numbered citations you can click. ChatGPT surfaces source links when a question triggers a live web search, so phrase it to require current information, then ask which sources it used. Log the domain behind each citation. After ten or twenty questions, the same handful of domains keeps reappearing, and that short list is your real target set. It is usually nothing like the list you would have guessed.
Why does Perplexity cite Reddit so heavily?+
Because it crawls the live web continuously and weights community-validated discussion, where the other engines lean differently. Reported figures vary a lot: one analysis put Reddit at roughly 46.7 percent of Perplexity's top citations, another tracking programme covering more than 200 million prompts put it at up to 20 percent of responses. Those numbers are not reconcilable, which is itself the useful lesson. The direction is consistent and the magnitude is contested.
Did ChatGPT stop citing Reddit?+
It appears to have dropped sharply. Practitioners tracking citation share reported Reddit holding around 3.8 percent of ChatGPT Search citations from mid-July into early August 2026, then falling to roughly 0.5 percent by 14 August, an approximately 86 percent relative decline in days. One operator running per-brand citation tracking reported Reddit citations at zero for a client in the current window, with the engine leaning on primary and manufacturer sources instead, while noting it was a single-account snapshot. We have not verified this independently and are reporting it as what practitioners observed.
How many sources does each engine cite per answer?+
Reported averages differ substantially by engine, with Perplexity around 21 sources per response against roughly 7 for ChatGPT in one analysis. The practical consequence matters more than the exact figures: an engine citing three times as many sources gives you three times as many ways in, but each individual citation carries proportionally less weight in the final answer. High citation counts are easier to enter and harder to dominate.
Should you optimise separately for each AI engine?+
Partly. The efficient approach is to treat the overlap set as your foundation and engine-specific sources as the layer on top. Domains repeatedly named as appearing across engines include Reddit, Wikipedia, YouTube, G2 and major business press, and presence in those has the broadest carry. Building separately for every engine is usually not affordable for a small team, and the divergence is large enough that ignoring it entirely leaves you invisible in whichever engine your buyers actually use.
How stable are these citation patterns?+
Not stable at all, and this is the most important caveat on the whole topic. The ChatGPT Reddit decline moved a major source from roughly 3.8 percent to near zero within about two weeks. Any published list of most-cited domains is a snapshot with a short shelf life, including the figures on this page. The durable practice is not memorising a source list, it is running your own measurement on a schedule so you see your category's list change when it changes.