Answer Engine Optimization · Measurement
AI Search ROI: The Attribution Problem Nobody Solves
Every AEO vendor eventually meets a CFO who wants a number. The honest answer is that classic channel ROI cannot be computed here, and the useful answer is what you can defend instead. This is both.
Key takeaways
- The classic ROI formula needs revenue attributable to a channel. AI search breaks that term: no click, stripped referrers, and influence that lands weeks before the session that converts.
- The recognisable signature is flat or falling organic traffic alongside steady pipeline. One marketer saw traffic down 12 to 15 percent while demo bookings held and customers volunteered that they had asked ChatGPT.
- Self-reported attribution at the point of conversion is the single most useful method available, because it is the only one that can capture a zero-click influence.
- Referrer-based tracking gives you a floor, never a total. Report it as a floor and say so.
- Dollar-per-mention frameworks are planning heuristics built from assumed click and conversion rates. Fine for sizing an opportunity, dangerous for reporting results.
01Why the normal formula breaks
Marketing ROI is normally simple arithmetic: take the revenue a channel produced, subtract what you spent, divide by the spend. The difficulty with AI search is entirely in the first term.
Three separate mechanisms break attribution, and they compound.
| Mechanism | What happens | What it does to measurement |
|---|---|---|
| Zero-click influence | The buyer reads an answer naming you and never clicks anything. | No session exists to attribute. The influence is invisible to analytics entirely. |
| Referrer inconsistency | Assistants differ in whether and how they pass a referrer. | Whatever you capture is a partial sample of the minority who clicked. |
| Delayed conversion | The answer is read weeks before the buyer acts. | The eventual session looks like direct or branded search. |
The third is the most underrated. Even a buyer who does click may not convert that day. When they return a fortnight later by typing your name, every analytics system on the market will credit direct traffic or branded search, and the AI answer that created the demand receives nothing.
02The signature: flat traffic, steady pipeline
There is a recognisable pattern that tells you this shift is happening to you, and it looks like bad news in a dashboard. A SaaS marketer described it precisely while trying to work out what to buy:
“Our organic traffic has been basically flat since around September, down maybe 12 to 15%, while demo bookings held steady. When I asked new customers how they found us, three in the last month said some version of 'we asked ChatGPT to compare options.' So the traffic didn't disappear, it moved somewhere I can't measure. Search Console shows nothing for this.”
Every element of that is diagnostic. Traffic down. Conversions flat. Customers volunteering an AI assistant when asked directly. Search Console showing nothing, because Search Console reports Google search and this did not happen in Google search.
Note how the evidence arrived: by asking customers. No tool surfaced it. That is not a gap in their stack, it is the structure of the problem, and it points directly at the most useful method available.
03What is actually measurable
Split the problem into three layers with honestly different confidence levels. Conflating them is what produces the confident numbers that later collapse.
| Layer | Confidence | What it tells you |
|---|---|---|
| Measured visibility | High. You control the measurement. | Whether you appear, and whether you are recommended or just mentioned. |
| Self-reported attribution | Medium. Real buyers, imperfect recall. | That AI search influenced specific closed deals. |
| Referrer-tracked sessions | High but partial. | A floor on arrivals. Never a total. |
| Branded search correlation | Low. Circumstantial. | Directional support, never proof. |
| Modelled dollar per mention | Assumption, not measurement. | Useful for sizing. Not for reporting results. |
The top layer is the one you fully control, and it is why measurement rather than optimisation is the defensible core of this discipline. The methodology is covered in measuring AI search visibility, and the distinction between being mentioned and being recommended matters more here than anywhere else. A programme that doubles mentions while recommendations stay flat has not moved revenue, and reporting a blended score would hide that.
04Self-reported attribution
One question on your demo form or in onboarding will tell you more about AI search influence than any analytics configuration. Ask how the person first came across you, and leave room for a free-text answer.
This works because it is the only method that survives a zero-click journey. The buyer read an answer, remembered a name, and arrived later by another route. No system observed that except the buyer.
Practical notes that make the difference between signal and noise:
- Ask how they first heard of you, not how they got to the site today. The second question collects the last click, which you already have.
- Leave it open-ended, or include an explicit option. A dropdown without an AI-assistant option will never record one, and you will conclude it is not happening.
- Ask at conversion, not in a later survey. Recall decays fast and response rates collapse after the fact.
- Log it against the closed deal, not just the lead. An answer attached to a lead tells you about interest. The same answer attached to closed revenue is what makes the case, and it costs nothing extra to carry the field through.
- Read the verbatims yourself. The phrasing tells you which assistant and often which question, which feeds directly back into your prompt set. See finding the prompts your buyers ask.
The known weaknesses are worth stating rather than hiding: people misremember, some skip the field, and a buyer influenced by an assistant six weeks ago may genuinely credit a colleague’s recommendation. This undercounts. It does not fabricate, which makes it the opposite failure mode to modelled attribution, and the safer one to build a case on.
05On putting a dollar value on a mention
Several vendors publish frameworks that assign a monetary value to an AI brand mention. Understand what those numbers are before you put one in a board deck.
The construction is generally the same: take an estimate of how many people saw an answer, apply an assumed click-through rate borrowed from traditional search, apply your conversion rate and average deal size, and multiply. The arithmetic is sound. Every input before the last two is an assumption, and the first one, how many people saw the answer, is not observable by anyone outside the assistant vendor.
That makes these frameworks genuinely useful for one job and dangerous for another:
| Use | Appropriate? | Why |
|---|---|---|
| Sizing whether to invest at all | Yes. | Order-of-magnitude reasoning is exactly what assumptions are for. |
| Comparing two channels before committing | With caution. | Only if you apply equally honest assumptions to the other channel. |
| Reporting what the programme delivered | No. | The output is generated by your inputs, so it will always confirm you. |
| Justifying renewal to a CFO | No. | A CFO who spots the assumption chain will discount everything else you said. |
The test is simple: if the number would change materially when somebody edits a cell containing an assumed rate, it is a model. Label it as one and it stays useful. Present it as a measurement and it costs you credibility the first time anyone asks how it was derived.
06What changes in buyer behaviour
The attribution problem exists because buying behaviour changed shape, not because analytics got worse. Understanding the new shape tells you which numbers are worth chasing and which are permanently gone.
In the search era, a buyer with a problem typed a query, scanned results, clicked several, and formed a shortlist from pages they visited. Every step after the query produced a measurable event on somebody’s property. The shortlist was assembled out of clicks.
In the assistant era, the shortlist is assembled inside the conversation. The buyer describes their constraint, receives three or four names with reasoning attached, and only then decides whose site is worth visiting. By the time you see a session, you have already won or lost the part that mattered.
| Stage | Search era | Assistant era |
|---|---|---|
| Problem framing | Compressed into a keyword. | Described in full, with constraints attached. |
| Shortlist formation | Built from pages the buyer clicked. | Built inside the answer, before any click. |
| Your first measurable event | The click that started the research. | A visit that happens after the shortlist exists. |
| What losing looks like | Ranking below a competitor. | Silence. You were never named. |
That last row is the one that should worry a marketing team most. Losing a ranking is visible in every tool you own. Not being named in an answer generates no data anywhere, produces no impression, and leaves no trace. You cannot detect it by watching your own properties, which is precisely why the visibility measurement in layer one is not optional.
There is one genuine upside hiding in this. Because the assistant reasons out loud, a measured answer tells you why a competitor was preferred, in plain language. Traditional rank tracking never offered that. You could see that you ranked fourth and never learn what the top result had that you lacked. That reasoning is the most actionable output in the whole discipline, and it is free once you are already measuring.
07What to tell a board
Boards handle uncertainty better than they handle being wrong. Report three things with their confidence attached.
- Measured visibility, stated as hard data. The share of buying questions where you appear, split into mentioned and recommended, per engine, on a frozen prompt set. This is the number you own and can defend line by line.
- Self-reported attribution, stated as direct evidence. The count of closed deals where the buyer named an AI assistant. Small numbers early. Real buyers, real deals.
- Branded search trend, stated as correlation. If branded search rises alongside measured visibility, say that and explicitly say it is circumstantial.
Two reporting mistakes are worth naming because both are common and both are self-inflicted. The first is leading with the modelled number because it is the biggest one on the page. It will be the first thing questioned and it will contaminate the credibility of the two real numbers sitting underneath it. The second is reporting visibility as a single blended percentage. A board hears one number, anchors on it, and then asks why it moved two points, which is a conversation about noise rather than about the business.
A more durable framing is to report the competitive position rather than your own score. Boards understand losing to a named rival far better than they understand a visibility index, and the comparison survives model updates that shift everyone’s absolute numbers at once. If every vendor in your category drops when an engine changes its retrieval, your score falls and your position does not, and only one of those two facts describes your business.
08How long it takes
The gap between doing the work and seeing revenue is longer than most programmes are funded for, and this is the practical reason AEO budgets get cut.
The chain is long: publish, get crawled, get incorporated into what an assistant retrieves, influence a buyer, and then have that buyer complete your normal sales cycle. Each step has its own lag, and the last one is entirely outside the programme’s control.
| Stage | Realistic lag | What to watch |
|---|---|---|
| Published to crawled | Days to weeks | Index status. |
| Crawled to cited in answers | Weeks | Citation appearances on your frozen prompt set. |
| Cited to recommended | Longer, and not guaranteed | The mentioned versus recommended split. |
| Recommended to pipeline | One full sales cycle | Self-reported attribution on new deals. |
The lag has one more consequence worth planning around: the work that eventually produces the revenue signal was usually done two or three months before anyone can see it, which means the programme most at risk of cancellation is the one that is about to start working. If you are going to cut it, cut it before you start rather than in month four, because a half-funded programme abandoned mid-lag produces the cost with none of the return.
For B2B software with a multi-month cycle, expecting a revenue signal inside one quarter is unrealistic, and promising one is how a programme gets cancelled in month four. Measure visibility monthly because it moves. Measure revenue effects across quarters because they do not.
The layer you can actually defend
Nothing on this page requires a tool. The self-reported attribution question is a form field, and the branded search trend is already in your analytics. Answer Radar covers the top layer: running a frozen prompt set on a schedule, recording mentioned, recommended and cited separately per engine, so the hard-data half of your board report is something you measured rather than modelled. We build in this category, so treat this as the disclosure it is.
09Frequently asked questions
Frequently asked questions
Can you measure the ROI of AI search visibility?+
Not in the classic sense, and any vendor promising a clean channel ROI figure is selling a model rather than a measurement. The standard formula needs revenue attributable to a channel, and AI search breaks that first term in three ways: much of the influence happens with no click at all, referrer data is frequently stripped or inconsistent across assistants, and the influence often lands weeks before the session that eventually converts. What you can measure is influence, isolated arrivals, and self-reported attribution. Those three together are a defensible case. They are not a channel ROI number.
Why does traffic stay flat while AI search is clearly working?+
Because AI search moves demand rather than creating a new traffic source. A SaaS marketer described exactly this: organic traffic down roughly 12 to 15 percent while demo bookings held steady, and three new customers in one month saying they had asked ChatGPT to compare options. Their conclusion was that the traffic did not disappear, it moved somewhere they could not measure, and Search Console showed nothing for it. Flat traffic with steady pipeline is the signature of this shift, not evidence that nothing is happening.
What is the single most useful AI search attribution method?+
Self-reported attribution, asked at the point of conversion. A free-text or single-select question on the demo form or during onboarding, asking how the person first came across you, captures the one thing no analytics tool can see. It is imperfect, because people misremember and some skip it, but it is the only method that captures a zero-click influence. Teams that add it typically discover AI-assistant mentions appearing in the answers within weeks, which is often the first hard evidence anyone has.
How do you isolate the traffic that does arrive from AI assistants?+
Segment on referrer where it exists and treat the result as a floor, never a total. Assistants that pass a referrer will show up in analytics as their own domain, and those sessions can be tracked normally. What you cannot capture is the far larger group who read an answer, did not click, and arrived later by typing your name into a browser. Those show up as direct or branded search. A rise in branded search that coincides with a rise in measured AI visibility is circumstantial evidence worth reporting as exactly that.
What should you report to a board about AI search?+
Report three things and be explicit about the confidence of each. Measured visibility, which is hard data you control: share of buying questions where you appear, and whether you are recommended or merely mentioned. Self-reported attribution, which is soft but direct evidence from real buyers. And a leading-indicator trend such as branded search volume, presented as correlation rather than proof. A board will accept an honest framework with stated uncertainty. It will not forgive a confident number that later turns out to be modelled.
Is a dollar value per AI brand mention credible?+
Treat it as a planning heuristic, not a measurement. Several vendors publish frameworks that assign a monetary value to an AI brand mention, typically by borrowing an assumed click-through and conversion rate from traditional search. The arithmetic is internally consistent and the inputs are assumptions. That is fine for sizing an opportunity and dangerous for reporting results, because the number is generated by your own assumptions rather than observed from your own funnel.
How long before AI search work shows up in revenue?+
Longer than most programmes are given, which is the practical reason AEO retainers get cancelled. Content has to be published, crawled, and then incorporated into what an assistant retrieves, and the buyer influenced by it still has to move through your normal sales cycle. For B2B software with a multi-month cycle, expecting revenue signal inside one quarter is unrealistic. Measure visibility monthly because it moves; measure revenue effects across quarters because they do not.