AI search and AEO
How to Test an AI Visibility Tool Before You Pay
Every buying thread names the same trackers and nobody has compared their numbers. Here is the acceptance test that replaces the comparison.
Key takeaways
- No independent accuracy comparison of these trackers exists, so the recommendations circulating in buying threads are secondhand.
- Run the same prompt five times before you trust a single number. If the spread between your runs is wider than the weekly movement on the dashboard, the dashboard is charting noise.
- Ask which surface the vendor samples. An API answer and a logged-in app answer are different products, and the gap between them is usually larger than the change you are being sold.
- A blended cross-engine score hides the fact you needed: which engine you are losing.
- Check that AI crawlers can reach your site before you buy anything. A blocked crawler invalidates every number that follows.
01What tools actually work for tracking how a brand is recommended across ChatGPT, Perplexity and Gemini?
Nobody has established which ones work, because no independent accuracy comparison of these trackers has been published. What works is the tool whose numbers you check against your own re-runs before the trial expires, and that takes an afternoon.
Read the buying threads and the evidence has an unmistakable shape. The same short list of vendors gets named, always secondhand.
“I've seen more than a few SEOs speak highly of Profound for LLM visibility tracking. Haven't used it myself yet though.”
That is not a criticism of the commenter, it is the state of the evidence. The most useful contribution in that thread was a map of the category, not a verdict.
“The leaders in the space focusing just on AI visibility seem to be: Profound, Searchify, Peec AI, AthenaHQ.”
The published guides do not fill it. The better listicles, such as Omniscient Digital on the five best AI visibility tools, grade on engine coverage, citation depth and pricing transparency. Those are purchase criteria, not accuracy criteria. They tell you what a tool watches, not whether what it reports is true. So build the test yourself: five checks, in this order, because a failure at any one invalidates the checks after it.
02Check one: does the number survive a re-run?
Pick one prompt where the tool says you appear. Run it yourself five times in one sitting, in a fresh session each time, and count how many of the five name your brand. That spread is the tool's margin of error, and you just measured it for free. This is the check practitioners arrive at first, and it reframes the category: answers are not rankings.
“manual prompt testing is a total trap. Since LLMs are non-deterministic, a single screenshot is basically useless data. You can ask ChatGPT the same thing twice and get two different answers”
The instability sits in the inference layer, not in your prompting. An agency teardown cites a lab study that sent 1,000 identical prompts at temperature zero, the setting meant to force deterministic output, and got 80 unique completions back. For the mechanism, Search Scope walks through why the same model does not repeat itself.
So the useful unit is persistence, not position: does the same question return you on day one, day four and day nine. One vendor has built exactly that and described it in public, which is a better signal than any feature list.
“One thing we've recently built into Matekio is Phrase Persistance insights. That is how frequently is a brand cited for a specific search phrase, overtime.”
03Check two: which surface did the tool sample?
Ask the vendor one question in the trial: which surface do you query, how often, and from where. Provider API, logged-out app, or a logged-in account. The answer changes what the number means more than any other variable.
The API is a different product from the app your buyer uses. It is stateless, carries no memory or custom instructions, and runs a system prompt you never see. The same teardown cites a February 2026 experiment of 900 trials across three OpenAI access surfaces in one day, where a single brand swung 32 percentage points between surfaces, per Search Scope's summary of that work. We have not reproduced those trials, so read them as one measurement rather than a settled constant.
The disclosure itself is the test. Practitioners already treat vagueness here as the tell.
“so if a tool can't clearly explain where its data comes from or how it's measured, that's usually a sign to be careful”
There is a fair defence of the vendors, and it is the more common position.
“their data isn't a 100% accurate since for an individual the AI answers would be somewhat personalized”
Both are true. A tool sampling one surface consistently is a fine directional instrument. Presenting it as a board number is the failure mode.
04Check three: do the engines disagree, and does the tool show you where?
Take one category prompt, run it in ChatGPT, Perplexity and Gemini yourself, and write down the brands each names. Then check whether the tool shows those lists separately or blends them into one score. The engines disagree, and the size of the gap is why a blended score is close to useless.
“I tested Chatgpt, Gemini and Perplexity with same set of 50 prompts, each AI search engine has different taste. All 3 AI engines agree on only 21% of the brands.”
That is one practitioner's test, not a published study, and we could not verify the prompt set behind it. It matches the direction of everything else in the threads.
“This is why traditional SEO metrics alone won't explain LLM visibility anymore. Different models clearly trust different authority signals.”
Watch the sources too. Buyers who check the links under an answer usually find it was assembled from third-party pages, not from any brand's own site.
“I always get recommendations of bigger brands on ChatGPT. When i check the links, it's using Amazon as a reference.”
| What the tool shows | What you can do with it | Verdict |
|---|---|---|
| One blended visibility score | Report a trend line upward | Fails, the engine you are losing is averaged away |
| Per-engine mention rate | See which engine to work on first | Minimum acceptable |
| Per-engine cited sources | See which third-party page to try to appear on | The version worth paying for |
| The full answer text, stored | Read how you were described, not just whether you appeared | The difference between a score and evidence |
05Check four: is the prompt set discovery-weighted or vanity-weighted?
Open the tool's default prompt list and count how many contain your brand name. If most do, the score will be high and meaningless: you are measuring whether an engine can look up a company it was just handed.
“the funny thing about ai visibility is everyone searches their own brand name, sees themselves mentioned, and celebrates. that's like searching your own house on google maps and claiming you've mastered local seo.”
The practitioner version of a good prompt set is specific about composition, not volume. One product marketer running this for a mid-market team keeps the operating view to five fields, and the first is the whole argument.
“Stable prompt set: product, category, competitor, and problem-aware prompts, not just brand terms.”
The rest of that sheet is worth copying: the engine result including who appeared instead of you, the class of source the answer was built from, an action owner and a review date. Four of the five fields are about acting on a loss and one is about detecting it. Most dashboards invert that ratio. For the longer method, see which prompts to track for AI visibility.
06Check five: can the engines reach your site at all?
Run this before the other four, because it is free and it invalidates everything downstream. If your robots rules or firewall block AI crawlers, no tracker's number tells you anything about a problem you can fix, and the category will sell you the score anyway.
“they have a tool that will check if your site can be reached by the AI indexing bots, as they're so often accidentally blocked”
The phrase that matters is accidentally blocked. It is rarely a decision anyone remembers: a bot filter, a CDN default, a user agent nobody allow-listed.
Two preconditions belong in the same pass. Third-party pages carry more citation weight than owned pages in most categories, so check whether you appear on the listicles and review pages in your space. And confirm the tool separates a mention from a citation. The criteria in RankDots' write-up on tracker accuracy are useful here, in particular the argument that a stored snapshot of the answer is what makes a citation auditable rather than asserted.
07What the acceptance test cannot tell you
It cannot tell you what to do when you lose. That is the honest hole in the category: every thread explains detection in detail, and not one describes a repeatable method for changing an answer once you know it is wrong. Treat any vendor claiming to close that loop as unproven until you have watched it work on your own prompts.
It also cannot settle whether this is worth measuring yet. The strongest counter here is that the whole exercise is premature.
“While LLMs are suggesting more, they don't (yet) represent a significant portion of search. Maybe 5-10%. Gemini is going to dominate anything in AI Answers on Google. Build for humans.”
The sharper version of that position skips visibility entirely and watches behaviour instead.
“But right now don't focus on visibility. Focus on the behavior of inbound traffic from those sources.”
Both are reasonable if your category has not moved. The test is your own pipeline, not a market statistic: are buyers arriving with a shortlist they did not get from you. If not, a spreadsheet and ten prompts will hold you another quarter.
One last limit. The test cannot make the number an outcome, because what moves it sits elsewhere.
“Content specificity and novelty, followed by brand sentiment and reputation on non-owned platforms, are the two factors that move the needle in my experience and testing.”
Which is why the test ends on evidence rather than features. Can this tool show you the answer text, the sources behind it, and the same measurement repeated often enough to trust the direction. If it can, the score is a byproduct and the evidence is the product. If it cannot, you are buying a chart. On the structural limits of what any tracker sees inside one engine, see what ChatGPT citation trackers can and cannot see.
Measure it the way you would audit it
Frequently asked questions
Which AI visibility tool is the most accurate?+
Nobody knows, because no independent head-to-head accuracy comparison has been published. Practitioners name Profound, Peec AI, Searchify and AthenaHQ, but the recommendations are almost always secondhand. The practical substitute is an acceptance test you run yourself during a trial: re-run stability, surface disclosure, cross-engine disagreement, prompt-set composition and crawler access.
How many times should the same prompt be run before the number means anything?+
Enough times that the movement you care about is larger than the spread between runs. A single sample per prompt per day is reading noise. Run one prompt five times in one sitting, count how many runs name your brand, and treat that spread as the tool's margin of error. A weekly change smaller than the spread is not evidence.
Why do ChatGPT, Perplexity and Gemini give different answers to the same question?+
They retrieve from different sources and weigh authority differently, so their recommendation sets only partly overlap. One practitioner who ran the same 50 prompts through three engines reported the three agreed on roughly one brand in five. A single blended visibility score averages away the thing you needed to see, which is the engine where you are actually losing.
Does it matter whether a tool measures the API or the consumer app?+
It is the single largest source of disagreement between a dashboard and what a buyer sees. The API is stateless, with no memory, no custom instructions and a different system prompt from the app your buyer uses. Ask which surface the vendor samples, how often, and from where. A vendor that will not answer is telling you something.
Can you track AI visibility without buying a tool?+
For under about ten prompts, yes. A spreadsheet, a fixed prompt list and a weekly run works, and it teaches you what the tools are doing. It stops working once you need repeated runs across several engines with someone reading every answer, which is where a tool starts to look reasonable against the hours.
What should you do when the tool says you are losing a prompt?+
Read the answer, not the score. Find the sources the engine cited to build it, and check whether they are pages you could realistically appear on. Detection is the easy half. Every thread asking this stalls at the same point, because nobody has a reliable method for changing an answer once they know it is wrong.
Are free AI visibility audits worth running?+
As a precondition check, yes. A free audit that tells you whether AI crawlers can reach your site is worth more than a paid score you cannot act on, because a blocked crawler invalidates every downstream number. As a purchase decision, no. A one-shot audit has the same single-sample problem as everything else here.