Demand intelligence
Manual Reddit Research: What Your Report Can Claim
A founder offered free hand-built reports and asked one question: are the findings accurate. Almost nobody answered it. Here is the answer.
Key takeaways
- A manual pass produces reliable verbatims, reliable thread locations and reliable competitor names. It does not produce reliable proportions, and most hand-built reports fail by claiming one.
- The sample is self-selected, concentrated in a small number of heavy contributors, and decaying as posts get deleted. None of that is fixed by reading harder.
- Automation improves recall and recency, not validity. It draws from the same population, so a share of voice computed over more threads is still a share of voice over the wrong denominator.
- The fix is a declared sample: what you searched, when, how many threads you read, and what you could not see. That paragraph is what turns an opinion into a report.
01Can manual research produce an accurate report?
Yes for the qualitative half and no for the quantitative half, and the split is sharper than either the tool vendors or the sceptics admit. A person reading threads by hand can establish, with high confidence, that a complaint exists, what words buyers use for it, which competitor gets named alongside it, and which specific conversations are worth a reply today. That same person cannot establish how common the complaint is, whether it is growing, or what share of a market holds it. The report is trustworthy right up to the first percentage sign.
This is not a theoretical concern. The question was put directly to a founder community: someone offered free thirty-day reports covering brand mentions, leads, competitor complaints, switching signals and the subreddits where those conversations happen, and said explicitly that they wanted feedback on whether the findings were accurate and useful. The thread drew eighty comments. In the sample we captured, almost every reply was a founder dropping a product name and a URL. The accuracy question went unanswered, which is itself the finding: the people who respond to an offer of free research are the people who want free research, not a cross section of anyone.
The two replies that were not product submissions were both about the business model rather than the method. One went straight past the report to the thing a founder actually wants.
“everyone is building simple scrapper tools for reddit analysis finding leads etc. What I would personally pay for is commission affiliate based marketplace for early B2B SaaS that can actually bring paying customers”
The other asked the question every reader of a hand-built report eventually asks, which is whether a person or a script produced it, and then answered its own question about whether any of it converts.
“Did you build a tool to do this? If you did it's validated. Seen a few people build this exact tool and put out a similar test comment. I'm not sure this route gets you customers but why not”
02What is actually wrong with the sample?
Three things, and they compound. None of them is about effort, which is the only problem the pages ranking for this question address.
The frame is undefined. A survey researcher starts by naming the population, then draws from it in a way they can describe. Manual forum research starts with a search box. Whatever the ranking algorithm surfaced for the phrases you happened to try, on the day you tried them, becomes the sample. You cannot state a response rate because there was no invitation, and you cannot state a margin of error because there was no draw. Compare that with how a research organisation does the same job: for its study of a large parenting forum, Pew Research Center ran an hourly collection pipeline for six months, captured every submission and its comments, and still paired the whole thing with a separate representative survey before drawing conclusions about parents.
Contribution is extremely concentrated. In that same six-month corpus, the top 5% of commenters produced around half, 52%, of all comments, per Pew Research Center. This is the old participation pattern that Nielsen Norman Group named the 90-9-1 rule: most users only read, a minority contribute occasionally, and a tiny group produces most of the content. Read a subreddit for a week and you are reading a few dozen people with unusual availability and unusual opinions. Their language is real. Their share of the market is not.
The evidence decays. Nine months after collection began, only 62% of the posts in the Pew dataset were still accessible: 29% had been deleted by their authors and 9% removed by moderators, according to the same study. A report that cites links rather than text will be partly unverifiable by the time someone acts on it, and the deletions are not random. People delete the posts they regret, which skews toward exactly the raw, specific complaints a competitor analysis values most.
03Which claims survive a hand-built sample?
Sort every sentence in the report into one of three tiers before you send it. The first tier is safe, the second needs a stated denominator, and the third should be deleted or demoted to a hypothesis with a test attached.
| Claim tier | Example | Manual verdict |
|---|---|---|
| Existence and language | Buyers describe this as seat pricing punishing part-time users | Safe. One clear thread proves the phrasing exists and it is quotable |
| Location | These conversations happen in three named communities, weekly | Safe if you say how you searched and over what window |
| Named alternatives | Two competitors get recommended in these threads | Safe as a list, unsafe as a ranking |
| Counts within your sample | Nine of the forty threads reviewed mention onboarding | Usable only with the denominator printed beside it |
| Proportions in the market | Most users of that tool are unhappy with support | Unsupportable. Delete it or restate it as a hypothesis |
| Trend | Complaints about pricing are increasing | Unsupportable from one pass. Needs the same query run twice |
The demotion in the last two rows is not pedantry. A founder who reads that most users of a rival are unhappy with support will build a campaign on it, lose, and never learn which link in the chain broke. A founder who reads that nine of forty reviewed threads named support, and that the forty were found through four specific searches in one month, knows precisely what they have: a lead worth testing with an actual survey or twenty sales calls.
Writing the tiers down also makes the report useful to somebody other than its author. Hand-built research usually lives in the head of the person who did it, and the tiers force the checkable parts to the surface.
04How do you make a manual report auditable?
Add one paragraph at the top that a stranger could use to reproduce your work. It is the difference between an opinion and a report. Six fields cover it.
- The queries. The exact phrases searched, not a description of them. Different phrasings surface different communities entirely.
- The window. The date range of the posts you accepted, and the date you collected. Both, because they are rarely the same.
- The volume. Threads opened, threads discarded, threads quoted. Three numbers, whose ratio tells the reader how much noise you waded through.
- The exclusions. What you skipped and why: promotional posts, threads with no replies, communities outside your market.
- The archive. The quoted text pasted into the report, not just the link, because a share of the sources will be gone.
- The blind spots. One honest sentence on what this pass could not see. Private communities, closed forums, everyone who never posts.
Nobody publishing about this workflow does that last one, and it is the field that earns the most trust. A report that says which questions it cannot answer is a report you can act on for the questions it does answer.
05When is a manual pass genuinely enough?
More often than the tool pages want to admit, and never for the deliverables they imply. What decides it is whether the output is a list of specific things to do or a number somebody will quote later.
| Job | Manual pass | Why |
|---|---|---|
| Find ten conversations worth replying to this week | Enough | Judgement per thread matters more than coverage |
| Collect the words buyers use for your problem | Enough | Saturation arrives fast and language does not need a denominator |
| Check whether anyone discusses a problem at all | Enough | One clear thread answers a yes or no question |
| Read the complaints about one named competitor | Enough, with the sample declared | A finite, searchable target and a stated window |
| Report share of voice against three rivals | Not enough | The denominator is unknowable by hand and the number will be quoted |
| Show whether sentiment moved this quarter | Not enough | Requires the identical query run on a schedule, twice |
| Catch switching conversations as they happen | Not enough | Threads go cold in a day and a person cannot watch continuously |
Notice that the first four are the jobs a small team actually has this week, which is why the honest verdict on manual research is far kinder than the category admits. What fails is the second half of the list, and it fails on structure, not diligence. For the broader method behind the first half, our guide to running market research on Reddit covers the search and saturation mechanics in more detail.
06Does automating it make the report more accurate?
It makes it more complete, more current and repeatable. Those are real gains and they are not accuracy. Automation improves recall, so you see threads your searches missed. It improves recency, so you reach a switching conversation while the person is still reading replies. It improves repeatability, so the same query run in October is comparable with the one run in September, which is the only honest route to a trend claim. What it does not do is change who was talking.
The population is identical. A crawler reads the same self-selected, concentrated, partly deleted set of posts that a person reads, and computing a percentage over ten thousand of them produces a confident number with the same bias underneath. Collection is also capped on both sides: Reddit listings return a maximum of 1000 items, returned 100 at a time, so an automated pass through the public API hits a ceiling too. Nobody has the full corpus.
This is where the pages currently ranking for the question go wrong. The strongest of them is a vendor use case page arguing that manual competitor monitoring does not scale, and its five supporting bullets are all about effort, consistency and timing. Every one of them is true. None of them is about whether the resulting report is correct. The category has quietly substituted the word scale for the word accurate, and a founder reading it comes away believing a subscription fixes a sampling problem.
The parts genuinely worth automating are the parts a person is bad at: watching continuously, catching a thread within hours, running the identical query on a schedule so two runs can be compared, and keeping a copy of text that later gets deleted. Judgement about which signals matter stays with the person either way.
07Complaint, or switching signal?
The failure mode that wastes the most time is treating every complaint as intent. Most venting is venting. Someone annoyed by a pricing change is not shopping; they are annoyed. The distinction is learnable and it is the part of the workflow that most rewards a human read.
Three markers separate the two. First, ownership: the person has the problem now rather than describing it in general. Second, movement: they mention evaluating, trialling, quoting, or asking for alternatives, not just disliking what they have. Third, reachability: enough context in the thread to say something specific and useful rather than a generic pitch. Threads that carry all three are rare, which is exactly why the count of them belongs in the report and the count of mentions does not.
The other reason to grade rather than count is that founders ask a locating question, not a volume question. One reply in the source thread put it precisely, describing the product first and then what they actually wanted to know.
“We're building for SaaS companies that want to turn their existing product into an agent that can actually complete workflows, not just expose APIs to an LLM. It would be interesting to see where founders and product teams are discussing agentic readiness, safe autonomous actions, and the workflows they would trust an agent to handle”
That is a where question, and a manual pass answers it well. It is when the same report starts answering how many that it quietly stops being evidence. For the grading criteria themselves, see our breakdown of switching intent signals.
One admission to close on. We could not verify one of the pages ranking for this question: it refused every fetch we attempted, so its argument is represented here only by its title, which promises to save hours. That is the shape of the whole SERP. Time is the problem everybody solves, and accuracy is the problem nobody states.
Keep the judgement, drop the collection
Frequently asked questions
Can a manual pass over forum threads produce an accurate report?+
It produces accurate quotes and accurate thread locations. It does not produce accurate proportions. A person searching by hand builds a sample they did not define, drawn from a population they cannot see, and any sentence in the report that starts with most people or the majority is a guess wearing a number. Keep the verbatims, cut the shares.
How many threads do you need to read before the findings are stable?+
There is no clean threshold. The practical stopping rule is saturation: read until three or four consecutive threads produce no phrasing, complaint or competitor you have not already recorded. If you hit that in twelve threads, the topic is narrow. If forty threads still produce new material, you are sampling a category rather than a problem, and the report should say so.
Why do the same complaints keep coming from the same handful of accounts?+
Because contribution is extremely concentrated. In a six-month study of one large forum, Pew Research Center found that a small fraction of commenters produced around half of all comments. Read a subreddit for a week and you are largely reading a few dozen people. That is fine for finding language and useless for estimating how common a problem is.
Does the evidence stay where you found it?+
Often not. In the same Pew dataset, only about six in ten collected posts were still accessible nine months later, and most of the losses came from authors deleting their own posts rather than from moderators. A report that links to threads without archiving the text will partially rot before anyone acts on it. Copy the quote into the report rather than citing a link.
Does a tool make the report more accurate than doing it by hand?+
It makes it more complete and more current, which is not the same thing. Automation improves recall, recency and repeatability. It draws from the same self-selected population, so it inherits the same bias, and a wrong proportion computed over more threads is still wrong. The accuracy gain comes from declaring the sample and grading each signal, not from the collection method.
When is a manual pass genuinely enough?+
When the output is a list of specific threads to act on, a set of phrases to use in copy, or a yes or no on whether a problem is discussed at all. One person, one afternoon, one competitor. It stops being enough the moment the deliverable needs a trend line, a share of voice, a rate of switching, or a claim that will outlive the week it was written.
What is the single most common error in these reports?+
Counting mentions and calling the count demand. Mention volume tracks how loud a community is about a topic, not how many buyers exist. The report should rank threads by whether the person owns the problem, is evaluating options now, and is reachable, then say plainly how many threads were reviewed to find them.