Answer Engine Optimization · Method

Buyer Prompts: You Cannot See What People Ask ChatGPT

Prompt data is not observable. OpenAI does not publish a query log and nobody can read private conversations, which means every prompt list on the market is an inference. That is fine, as long as you know it. Here is how to build one that holds up.

By Linkeddit·Updated August 28, 2026·15 min read

Key takeaways

  • There is no prompt equivalent of keyword volume. ChatGPT conversations are private and no public query log exists, so every prompt list you can buy is an inference rather than a measurement.
  • That does not make prompt tracking useless. It makes the sourcing method the thing you should evaluate, and most vendors will not describe theirs.
  • Rewriting keywords as questions produces a set that looks right and behaves nothing like buyer language, because it keeps the keyword compression and adds a question mark.
  • One prompt is an anecdote. Answers vary run to run, so you read results at the topic level across many prompts, never at the individual prompt level.
  • Being mentioned in an answer is not being recommended. Tools that report one blended visibility score make those two situations indistinguishable.

01What buyer prompt discovery is

Buyer prompt discovery is the work of figuring out which questions real buyers ask AI assistants about your category, so you can track whether you show up in the answers. It is the input step for everything else in answer engine optimization. Get the prompt set wrong and every metric downstream is measuring the wrong thing precisely.

The problem is recognisable long before anyone has a name for it. A SaaS marketer described the shape of it exactly:

Our organic traffic has been basically flat since around September, down maybe 12 to 15%, while demo bookings held steady. When I asked new customers how they found us, three in the last month said some version of 'we asked ChatGPT to compare options.' So the traffic didn't disappear, it moved somewhere I can't measure. Search Console shows nothing for this and me manually typing prompts into ChatGPT every monday isn't a system.
SaaS marketer, via r/SEO_tools_reviews

That last clause is the honest statement of the problem. Manually typing prompts is not a system. But before you can build the system, you have to solve the sourcing question, and that is where the category gets quiet.

02The part nobody says out loud

You cannot see the prompts people actually type into ChatGPT. Nobody can. Conversations are private. OpenAI does not publish a query log. There is no Search Console for ChatGPT, no keyword planner, no volume figure. The observability that traditional SEO takes for granted simply does not exist here.

This matters because a large part of the tooling market is sold in language that implies otherwise. When a product offers you the questions your buyers ask ChatGPT, what it is actually offering is one of three things: public language harvested from forums, search data, and communities; aggregate patterns from its own customer base; or a language model’s guess at how a person might phrase the question. All three are inferences. Some are good inferences. None are measurements.

Saying this plainly costs nothing and buys a lot. It sets the right expectation for what a prompt set can tell you: it is a representative sample of plausible buyer language, useful for detecting change over time, and not a census of demand. Once you hold it that way, the rest of the method follows.

03Why a keyword list is not a prompt set

The most common shortcut is to take an existing keyword list and add question marks. It produces a set that looks reasonable and behaves nothing like real buyer language.

Keywords are compressed. Years of search behaviour trained people to strip context out, because search engines rewarded terse queries. “Best CRM” carries no budget, no team size, no existing stack, no constraint. Conversation removes that pressure, so people put the context back in. The realistic version is closer to “what CRM should a 12-person agency use if we already run everything through Google Workspace and we do not want another per-seat contract.”

KeywordThe prompt a real buyer asksWhat the difference changes
best CRMWhat CRM should a 12-person agency use if we already run on Google Workspace?Integration constraint becomes the deciding factor, so integration pages get cited.
competitive intelligence softwareHow do I keep track of competitors without paying enterprise prices?The answer set becomes budget tools, not the category leaders.
AI visibility toolHow do I check whether ChatGPT recommends us instead of our competitor?Shifts from product listings to method content, which changes who gets cited entirely.

The practical consequence: prompts that carry a constraint retrieve a different set of sources than the bare category term. If your prompt set is all bare category terms, you are measuring the head of the category and missing every answer where a specific constraint decides the recommendation. That is usually where the winnable ground is.

04The five prompt types worth tracking

A prompt set that only contains category questions will under-represent how buyers actually move through a decision. A practitioner working through this problem laid out a taxonomy that holds up well, and the framing at the end is the important part.

Prompt typeExample shapeWhat it tells you
CategoryWhat are the best platforms for tracking brand visibility?Whether you are in the consideration set at all.
ProblemHow can I measure whether my brand appears in AI search?Whether your method content gets cited before any product is named.
ComparisonBrand A vs Brand B for enterprise SEOHow you are characterised against a named rival, including inherited framing.
Use caseBest SEO platform for an international marketing teamWhether you win on specific constraints even when you lose the head term.
RecommendationWhat software should a B2B company use for AI search optimization?Whether you are the answer, rather than one of several names listed.
The important distinction is that these are buying and research questions, not simply keywords rewritten as prompts.
Practitioner, via r/SEMrush community

Problem prompts are the most under-used of the five. They fire early, before a buyer knows which products exist, and the sources cited in those answers shape the consideration set the buyer arrives with. If you only track category and comparison prompts, you are measuring the end of the funnel and ignoring the part you can most easily influence.

05Five sources you can actually observe

Since prompt data is not observable directly, you triangulate. Each of these sources is imperfect alone. A prompt that appears in three of the five is worth tracking.

  1. Recorded sales calls and discovery notes. The highest quality source available, and the most neglected. When a prospect says how they came across you or what they compared, that is real buyer language from a real buyer. Ask the question explicitly on every discovery call and log the verbatim answer.
  2. Your own Search Console long tail. Filter to queries of eight or more words. Conversational, constraint-carrying queries are already arriving in traditional search and they are the closest observable proxy for prompt phrasing you own outright.
  3. Support tickets and onboarding questions. The questions customers ask in week one are the questions prospects ask an assistant in week zero. This source is free, sitting in your helpdesk, and almost nobody mines it.
  4. Public community discussion. Threads where someone asks for a recommendation in your category are buyer prompts typed in public. The phrasing is unpolished and constraint-heavy, which is exactly what makes it useful.
  5. The assistants themselves. Ask an assistant what related questions people commonly ask about your category. It is an inference, and you should label it as one, but it is a cheap way to expand coverage once the other four have given you a spine.

06One prompt is an anecdote, not data

AI answers vary between runs even when the prompt is identical. This means a single prompt result carries almost no information, and a prompt set organised as a flat list of hundreds of individual prompts produces volume without direction.

A founder who had run dozens of demo calls walking people through prompt selection described the mistake that ruins most self-built sets, and the fix:

Topics work like SEO clusters, one topic equals one cluster, and you read your analytics at the topic level. One prompt is an anecdote, not data.
AI visibility tool founder, via Reddit

The same practitioner flagged a second structural error: mixing brand names into a topic cluster. If a topic is meant to measure “best CRM for small agencies” and half the prompts inside it name specific vendors, the cluster is measuring two different things at once and the aggregate is meaningless.

Topic
The level you read results at
15 to 25
Prompts per topic for a usable distribution
Never
Times to change wording mid-series

The arithmetic matters more than it first appears. Two hundred prompts spread across forty topics gives you five observations per topic, which is noise. The same two hundred across ten topics gives you twenty per topic, which is a distribution you can actually read a trend from. Fewer topics tracked properly beats broad coverage tracked badly.

07Mentioned is not recommended

The most common measurement error in this category is treating appearance in an answer as a binary. It is not. Being named once inside a long answer is a fundamentally different outcome from being the recommendation the buyer acts on.

'Did the brand appear' is not a yes/no. Getting name-dropped in a wall of text is not the same as being the recommendation. I track those as two separate metrics now, one for 'mentioned at all', one for 'actually recommended'. Loads of brands score decent on the first.
Practitioner who built their own tracker, via r/GEO_optimization

Record at minimum three things per run, not one:

  • Mentioned. Did the brand name appear anywhere in the answer?
  • Recommended. Was it presented as a suggested option rather than listed in passing?
  • Cited. Was one of your pages used as a source, and which one? This is the only one of the three that gives you a direct action.

The citation field is the most operationally useful because it turns a score into a task list. Knowing you are invisible is not actionable. Knowing that a specific competitor page is being cited for a topic you also cover is a piece of work you can assign this week.

One further caution on engine differences. Vendors in this space commonly claim that different assistants lean on structurally different source types. That claim is plausible and widely repeated, but we have not verified it against a controlled test, and you should treat any per-engine breakdown as a vendor claim until you have run your own prompts and looked at the citations yourself.

08Cadence, drift, and version control

A prompt set is only useful if it stays comparable across runs, which makes version control more important than coverage. The failure mode is quiet: someone improves the wording of a few prompts between runs, the numbers move, and the team reads that movement as a change in visibility when it is a change in the instrument.

The manual version of this problem is well described:

Every morning I manually copy the same prompts into ChatGPT and a few other tools to check whether our brand gets mentioned. I lose track of which prompts I ran, the answers change depending on wording, and comparing results week to week is messy.
Startup content lead, via r/SEO_tools_reviews

Four rules make a prompt set durable:

RuleWhyWhat breaks without it
Freeze wording within a seriesPhrasing changes retrieval.Your trend line measures your editing.
Version the set, do not edit itNew prompts start a new series.Old and new results silently average together.
Date and label every runModel updates land without notice.You cannot separate a model change from a real change.
Record the engine per resultEngines answer differently.A blended number hides which engine moved.

On cadence: monthly suits most B2B categories. Weekly runs mostly surface run-to-run variance and encourage over-reaction. Quarterly lets a shift settle in before you see it. If you only have budget or patience for one thing, run a smaller set monthly rather than a large set once.

For how to choose which topics to cover once you have the raw prompt pool, see which prompts to track for AI visibility, which covers the selection step this page feeds.

When the sourcing step is the bottleneck

The triangulation method above works with a spreadsheet, your helpdesk, and your Search Console, and you should run it that way first. Answer Radar exists for the part that does not survive manual effort: running a frozen, versioned prompt set on a schedule, recording mentioned, recommended, and cited separately per engine, and keeping runs comparable across model updates. We build in this category, so treat this as the disclosure it is.

See how Answer Radar works

09Frequently asked questions

Frequently asked questions

Can you see the actual prompts people type into ChatGPT?+

No. ChatGPT conversations are private and OpenAI does not publish a query log the way Google publishes Search Console data for your own site. There is no prompt equivalent of keyword volume. Any tool advertising real buyer prompts is giving you an inference built from public language, its own customer set, or a model's guess at how people phrase things. That inference can still be useful, but it is a model of demand rather than a measurement of it, and you should price your confidence accordingly.

How do you find the prompts your buyers ask about your category?+

Triangulate from sources you can observe. The five reliable ones are: recorded sales calls and discovery notes, where prospects describe how they searched; your own Search Console queries, especially the long conversational ones; support tickets and onboarding questions; public community discussion where people ask for recommendations in your category; and the assistants themselves, which will readily tell you what related questions people ask. No single source is sufficient. A prompt that appears in three of the five is worth tracking.

What is the difference between a keyword and a prompt?+

A keyword is a compressed query typed to a search index. A prompt is a full question asked in conversation, usually carrying context the keyword strips out: budget, team size, industry, and the constraint the person actually cares about. "Best CRM" is a keyword. "What CRM should a 12-person agency use if we already run everything through Google Workspace" is a prompt. Rewriting your keyword list as questions produces a prompt set that looks right and behaves nothing like real buyer language, because it keeps the compression and adds a question mark.

How many prompts should you track?+

Enough per topic that the result is a distribution rather than a coin flip. AI answers vary between runs even when the prompt is identical, so a single prompt tells you almost nothing. Practitioners running this at scale organise prompts into topic clusters and read results at the topic level rather than the prompt level. Tracking 200 prompts spread across 40 topics is weaker than tracking 200 prompts across 10 topics, because the second gives you 20 observations per topic instead of 5.

What types of buyer prompts should a prompt set include?+

Five types cover most commercial intent: category prompts ("what are the best platforms for X"), problem prompts ("how do I measure whether X is happening"), comparison prompts ("A vs B for enterprise"), use-case prompts ("best X for an international team"), and recommendation prompts ("what should a B2B company use for X"). The distinction that matters is that these are buying and research questions, not keywords rewritten as questions.

Is being mentioned in an AI answer the same as being recommended?+

No, and conflating them is the most common measurement error in this category. Being name-dropped inside a long answer is different from being the answer. A practitioner who built their own tracker described splitting this into two separate metrics after finding that many brands score reasonably on mention rate while almost never being the actual recommendation. If your tool reports one blended visibility number, you cannot tell which of those two situations you are in.

How often should you re-run a prompt set?+

Monthly is a reasonable default for most B2B categories, with the caveat that you must hold the prompt wording fixed between runs. Answers shift with model updates, with retrieval, and with small changes in phrasing, so a set that gets reworded each month produces a trend line that measures your editing rather than your visibility. Version the prompt set, date every run, and record which engine and which run produced each result.