Reddit is where people say what they actually think about your product, usually in a thread you were not invited to. It is the highest-candour feedback source on the open internet and simultaneously the easiest to misread, because everything you see has been through a voting filter before it reaches your eyes. Read a thread top-to-bottom and you have not measured opinion — you have measured what the ranking algorithm chose to show you, filtered again by how far you happened to scroll.
This guide is a framework that fits in an hour and survives being shown to someone else: define the question, export, correct for the sort, count themes weighted by score, check who is talking, and write a decision. It works for a founder researching a market, a brand auditing a launch, and an agency building a positioning brief.
What makes Reddit different from every other comment section
Four structural properties change how the analysis has to work.
- Votes, not just reactions. Score is net: upvotes minus downvotes. A comment at zero might have been ignored, or it might have been fought over by two hundred people. Both look identical in the number.
- Ranking with a strong early-vote bias. Reddit's own comment sorting documentation describes several orderings, and the default blends score with confidence and timing. The practical effect is well known to anyone who has commented late on a big thread: a good comment posted three hours in is buried under a mediocre one posted three minutes in.
- Threading that carries meaning. Elsewhere replies are a side conversation. On Reddit the reply chain is where the claim in the parent comment gets tested. The depth of a chain is itself a signal — five levels deep means something was contested.
- Community context. The same product gets a different reception in an enthusiast subreddit and a general one, and neither is wrong. A subreddit is a population, not a sample of the market.
Every one of these is invisible while scrolling and obvious in a spreadsheet, which is the entire argument for exporting before analysing.
Step 1: Write down the question before you look
The discipline that separates analysis from browsing is deciding, in one sentence and in advance, what decision this is supposed to inform. Good versions look like:
- "Which objection should the pricing page answer first?"
- "Is the complaint about onboarding a vocal minority or the median experience?"
- "Which competitor do people actually name when they say they switched?"
Bad versions look like "understand what people think about us". That produces a word cloud, and nobody has ever made a decision from a word cloud. Write the question at the top of the sheet. If a finding does not bear on it, note it and move on.
Step 2: Export the thread
Paste the thread URL into the Reddit comment exporter and run it with replies enabled. What comes back is one flat file with a consistent set of columns — author, username, text, likes (the score), replies, created_at, language, is_pinned, plus id and reply_to_id for reconstructing the tree. The mechanics, including how to handle very large threads, are in exporting Reddit thread comments.
Two notes on scope. If you are researching a topic rather than a single thread, export the three or four threads that matter and stack them in one sheet with a thread column — cross-thread patterns are far more trustworthy than anything inside one discussion. And if the question is about a specific product's reputation, search for threads where it is mentioned rather than only threads that are about it; incidental mentions in a "what do you all use?" thread are less performative and often more honest.
A word on the official API: Reddit's Data API terms introduced paid access tiers in 2023, and the historical archives many analysts relied on were shut down at the same time. That is worth knowing for two reasons: the free bulk-history era is over, and any old tutorial you find that pipes a public archive into a notebook no longer describes something that works.
Step 3: Correct for the sort before you count anything
This is the step nobody does, and it is where most Reddit analysis goes wrong.
Do two sorts in the sheet and compare them. First by likes descending — the consensus view, what the thread collectively endorsed. Then by created_at ascending — the chronological view, which reveals how the conversation actually developed. The gap between them is informative on its own. A theme that appears early and scores well was the thread's frame. A theme that appears repeatedly late with low scores is usually people arriving with a genuine problem after the thread's attention had moved on; low score there means low visibility, not low validity.
Then look explicitly at the bottom. Filter for comments with a score at or below zero and read them. Some are trolling and some are noise, but a systematic pattern of downvoted comments all making the same point is one of the most useful findings a Reddit analysis produces: it is a position the community rejects socially but which several people independently hold. If those people are your customers, that gap between what is safe to say in the subreddit and what people actually think is the finding.
Step 4: Count themes, weighted by score
Read the top hundred comments by score and write down the recurring themes as you go. Most threads collapse into five to eight. Tag each row with one, then count.
Count two numbers per theme, not one: how many comments carry it, and the total score across them. They diverge more than you would expect, and the divergence is the point. Twelve comments totalling 4,000 points is a view a large silent audience endorsed. Ninety comments totalling 300 points is a view many people hold but few others cared to affirm. Both are real; they are not the same finding and should not be reported the same way.
Three rules keep the counting honest:
- Tag once. A comment can touch three themes, but if you tag it three times your counts inflate toward whichever comments are longest. Tag the primary point.
- Tag before you know the totals. Once you have seen the counts you will unconsciously tag ambiguous comments toward the leading theme.
- Keep a "one-off" bucket. Singleton comments are not noise to delete — the same singleton appearing in three separate threads is an emerging theme, and you will only notice if you kept it.
Step 5: Score sentiment inside each theme, never across the thread
An overall sentiment score for a Reddit thread is close to meaningless. Threads are not homogeneous — they are several arguments occupying the same page, and averaging them produces a number that describes none of them. "62% positive" hides the one theme that is 95% negative and happens to be about your pricing.
Score inside themes instead, and use three buckets rather than a scale: positive, negative, mixed or asking. That last bucket matters more on Reddit than anywhere else, because a large share of the highest-value comments are neither praise nor complaint — they are people evaluating options, and their questions are a direct map of what your positioning fails to answer.
If you automate the scoring, know what you are buying. Sentiment models are trained largely on review-style text and degrade sharply on irony, sarcasm and community in-jokes — the persistent difficulty of automated sarcasm detection is a well-documented problem in the NLP literature, with dedicated shared tasks at venues like the ACL Anthology devoted to it. Reddit is arguably the worst-case input: "yeah, it's great" and "yeah, it's great 🙃" are opposite statements to a human and identical to most classifiers. Use automated scoring to triage and rank, then spot-check the top of every bucket by hand before any number leaves your desk. The sentiment analysis caveats that apply on other platforms apply doubly here.
Step 6: Check who is talking
Pivot on username. Two questions matter.
How many distinct people are in this thread? Four hundred comments from ninety accounts is a discussion. Four hundred comments from twenty-five accounts is a handful of people arguing at length, and reporting it as "400 comments of feedback" would be wrong. This one pivot has killed more overconfident Reddit findings than any other check.
Who is doing the heavy lifting? The accounts with the most comments in a thread are usually either domain experts answering everyone or the most invested critics. Read their comments as a set. An expert who patiently answers thirty questions is telling you what your documentation fails to explain; a critic writing eight long comments is giving you the fullest articulation of the objection you will ever get for free.
Then check the deep chains. Sort by reply_to_id and find the branches that go four or five levels deep. Depth means contested, and contested means the claim at the top was worth arguing about. Threads' most useful content is disproportionately buried five replies down where nobody scrolls — including, on the day they looked, you.
Worth naming: vote counts are not a clean measure of people. Reddit's own transparency reporting documents substantial ongoing enforcement against spam and manipulation, and displayed scores are deliberately fuzzed. Treat score as an ordering signal, not a precise headcount.
Step 7: Write the decision, not the report
Finish with three lines, each carrying its own evidence:
Theme 1 — "the free tier is too limited to evaluate" (34 comments, 2,900 points, negative). Highest score-per-comment in the thread; appears in all three threads exported. Raise the trial limit or add a sandbox — this is the top-of-funnel blocker.
Theme 2 — "how does it compare to [competitor]" (61 comments, 800 points, asking). The most common comment type and almost entirely unanswered. Build the comparison page; it is being asked for explicitly.
Theme 3 — "the mobile app crashes on import" (9 comments, −40 points combined, negative). Downvoted, so nearly invisible in the thread, but nine independent reports of the same bug. Score is a popularity measure, not a truth measure. File it.
Notice what is absent: an overall sentiment percentage, a word cloud, and a chart of comment volume with no interpretation. Those are artefacts of tooling, not findings. Everything above is something a team can act on this week, with a number attached to justify doing it.
Mistakes that make Reddit analysis worse than no analysis
- Treating one subreddit as the market. The enthusiast subreddit for your category is a population of people who care intensely about the category. Their priorities are systematically not the median buyer's. Their objections are still real; their weighting is not transferable.
- Reading only the top comments. That is reading the ranking, not the thread. Everything in step 3 exists to prevent this.
- Quoting a score as a headcount. Scores are fuzzed and net. A +200 comment did not get 200 people's approval.
- Ignoring thread age. A thread from eighteen months ago describes a product that no longer exists — yours or a competitor's. Sort your exported threads by date and weight recent ones heavily.
- Responding as a brand without thinking. Analysis is not engagement. Marketing-voiced replies in a critical thread reliably make things worse; if you respond at all, respond as a named person, answer the specific technical objection, and do not pitch.
Doing this repeatedly
Comment analysis compounds. The second thread you analyse is worth more than the first, because you can compare — and by the tenth you have a baseline that makes anomalies obvious without any modelling at all.
A workable cadence: export every thread that mentions your product or category monthly, append to one master sheet with thread and subreddit columns, and re-run the theme counts. Themes that persist across threads and across subreddits are structural — they are about your product, not about one discussion. Themes confined to a single thread usually belong to that thread's particular argument.
Because the export columns are identical across Reddit, TikTok, Instagram and YouTube, one master sheet can hold all of them without reshaping anything. That cross-platform view is where the framework pays off most: the objection that shows up in a Reddit thread, a TikTok comment section and a YouTube review within the same month is not an opinion, it is a product problem.
One last habit: keep the raw exports. Threads get locked, accounts delete their history, and moderators remove comments months after the fact. The file you pulled today is the only complete version of that thread that will still exist next year, and a copy in your own drive costs nothing.
Related reading
- How to export Reddit thread comments — the mechanics behind step 2
- Scraping Reddit comments — API, PRAW and the limits of each
- Exporting Reddit comments to Excel
- The same framework applied to YouTube
- Building a social listening report from real comments
- Reddit Comment Exporter — 3 free exports a day, no signup
