Every list gives you questions and stops there. This one pairs each question with the answer at scale, drawn from a million-plus unprompted accounts of the problems people already have.
There is no shortage of customer discovery question lists. We pulled the pages currently ranking for this question and they are remarkably similar to each other: between twelve and fifty questions, most of them good, nearly all traceable to the same three sources, and every one of them stops at the question mark. You are handed a list and left to find out for yourself what a useful answer sounds like, how to tell a real problem from a polite one, and how to get anyone on a call at all.
This page does the second half. Each question below is paired with what the answer looks like when you collect it at scale, drawn from the 1M+ documented complaints, reviews and forum accounts in our pain points database. Not because the data replaces talking to people, but because it tells you what to listen for, and because it lets you start today rather than in three weeks when the first call finally lands.
The standard failure is not that founders ask too few questions. It is that they ask questions whose answers cannot be wrong. If you describe your idea and ask whether it sounds useful, the person in front of you is being asked to evaluate something you obviously care about, in real time, at no cost to themselves. Agreement is the socially cheapest response, and you will get it.
The consequence is well documented at the other end of the funnel. CB Insights finds no market need is the most common cause of startup failure, appearing in 42% of post-mortems. Many of those founders did talk to users. They asked questions that could only return yes.
The second failure is subtler and harder to fix. Interviews are scheduled with people who are easy to schedule with, which means friends, former colleagues and the members of your own professional community. That group is systematically more like you than the market is, and it is systematically more encouraging. You can ask perfect questions and still get a distorted answer if the sample is drawn from the people who like you.
Discovery has two modes. The first is the one everybody writes about: you talk to people. The second is the one almost nobody treats as discovery: you read what people already wrote when no founder was present.
The second mode has a property the first cannot match. When someone writes a two-star review at eleven at night, they are not managing your feelings and they are not predicting their behaviour. They are describing what happened, while annoyed, in their own words, with the detail that annoyance produces. That is a better interview than most interviews, and there are millions of them already sitting in public.
The right sequence uses both. Read first, to learn the failure modes and the vocabulary. Interview second, to test the specific version that applies to your segment and to ask the follow-ups that only a live conversation permits. Founders who do it in that order arrive at the call already fluent and spend the thirty minutes on substance rather than on orientation.
These six establish whether a problem exists, how often, and how much of the person's attention it takes. Every one is past tense.
The answer you want has a date, a sequence and other people in it. The answer you do not want is a general statement about how things are. Here is what a good one looks like, written unprompted:
"We had a 250% spike in tickets due to a migration, now sitting on a 350-ticket backlog with no added headcount and the team is burning out." – r/managers
"I was told I would manage a new hire, then told I was not, then told I was. But I am not really being kept in the loop about everything he is working on." – r/managers
Both have a triggering event, a consequence and named participants. Neither is a hypothetical. This is the texture you are listening for, and once you have read a hundred of them you recognise it instantly.
"I was presented a plan and then fired within a week without completing it. I got one three weeks into my orientation." – r/careeradvice
"I deposited funds into my account. The transaction completed but the money still has not been credited after 42 days. Multiple reach outs but only vague replies." – r/trading
Forty-two days is the kind of detail that only appears when somebody is recounting something that actually happened to them. No survey question produces that number. Our writing on problems worth solving, business pain points and ideas backed by real complaints expands on the pattern.
If you only get to ask four questions, ask these. A workaround is proof that the problem is severe enough to spend something on, and the units the person uses to describe it tell you what they think it costs.
The fourth is the most underused question in customer discovery. If somebody is being paid to do a thing by hand, the budget exists, it has already survived a manager's approval, and it is measurable.
"I cut weekly ad reporting from 4 min to 30 sec. Every week your manager or client asks how CPC and CPA are trending across Google, Meta, LinkedIn, Reddit. Then you open 12 tabs, export CSVs, and wrestle VLOOKUPs." – r/remotejobs
"Monitoring performance metrics, comparing channel effectiveness, and generating client reports was an absolute nightmare of scattered data." – r/automation
"We need a hands-free experience for my team and cannot afford additional staff. It is not easy to configure, so we need to hire an AI specialist just for configuration and maintenance purposes." – r/automation
That last quote is the strongest possible answer to question ten. A business is describing hiring a specialist to compensate for software that does not configure itself. There is no ambiguity about whether a budget exists.
"I need something that feels like a product customizer, but with minimum order quantity and volume pricing baked in. Getting all of that working together usually means mixing several apps with some custom code." – r/wix
"I am working lean, trying to avoid using many tools and ending up with an unauditable or non-reproducible evidence base. I would really value input on structuring a review that stands up to regulatory scrutiny." – r/medtech
"After looking at over 100,000 job descriptions I found specific phrases that reliably indicate a company is actively evaluating solutions." – r/salesdevelopment
Notice what all three have in common. Each person built something by hand rather than buying it, and each can describe the construction in detail. When an interview produces that, you have found the product shape without having to guess at it. This is the same signal we use to surface single-feature micro SaaS ideas and underserved industries.
The same pattern shows up across freelance demand, which is effectively a market in workarounds. The recurring themes in the freelance job pain points we track are time-consuming manual rendering, error-prone lead generation, inefficient appointment scheduling, manual product listing management, error-prone bookkeeping and reconciliation, inefficient inventory management and manual proofreading. Almost every entry contains the word manual or the phrase error-prone. Our piece on validating SaaS demand with freelance jobs covers how to use that, with the caveat that the rate fields in that dataset are unpopulated so we never quote dollar figures from it.
Never ask how much someone would pay. Ask what they already pay, and for what, and who signed it off.
The fifteenth question is the one that separates a nice-to-have from a purchase. Every purchase has a trigger, and the trigger is usually an incident rather than a realisation. If nobody can name the incident, you are looking at a preference.
"I have heard how great this can be and want to try it out. Is there a realistic expectation of how much I should plan on spending for marketing and ad campaigns each month?" – r/dropshipping
People asking publicly what they should expect to spend are telling you two things at once: that a budget is being formed, and that no trustworthy benchmark exists for them to consult. Both are useful. Our SaaS metrics benchmarks and indie SaaS revenue data exist because that gap keeps showing up.
If the market already has products in it, the most valuable thing you can learn is why people leave them. These four do that.
Question nineteen is the closest thing to a legitimate hypothetical, and it works because people who have already switched once can answer it from experience rather than from imagination.
"It has been ok for my store, but I am trying to add certain products and they claim the products are restricted." – r/dropshipping
"Really good underlying model, but the app makes for very poor experience. Weird screen transitions, an annoying spinner on every screen, random vibrations." – App Store review
The first is a switching trigger in progress: a specific blocked task against an otherwise acceptable tool. The second is a complaint that almost never causes a switch on its own, which is exactly why the distinction matters. Interviewing people who are mid-switch is worth more than interviewing the merely dissatisfied, and why SaaS customers churn covers how to tell the two apart.
These five end the conversation in a way that produces either a next step or a clean negative, both of which are useful.
Question twenty-two is the only reliable way to escape your own network, and it compounds. Question twenty-three converts stated interest into a small commitment, which is the cheapest available lie detector. Question twenty-four regularly produces the single most useful sentence of the call.
| Do not ask | Why it fails | Ask instead |
|---|---|---|
| Would you use this? | Asks for a prediction, invites politeness | What are you using today, and what broke last time? |
| Do you think this is a good idea? | Asks for approval of you, not of the product | Who else have you seen try to solve this? |
| How much would you pay? | Nobody knows, and the answer is always too high | What do you currently spend on this? |
| Would you buy it if it did X? | Conditional futures cost nothing to agree with | What made you finally pay for the last tool you bought? |
| Is this a big problem for you? | Invites the answer that keeps the conversation pleasant | Is it in your top three, or further down? |
| What features would you want? | Turns the user into a designer | What did you have to do by hand last week? |
The best-known heuristic in this field is The Mom Test by Rob Fitzpatrick: a good question is one your own mother could not lie to you about. Do not pitch, do not hypothesise, ask about specific past behaviour. Almost every question above is an application of it, and if you read one thing on interview technique, read that.
Its limit is not conceptual, it is logistical. The Mom Test makes your questions better. It does not get anyone on the call. In practice the modal customer discovery effort does not fail because the questions were leading. It fails because eleven emails went out, one person replied, the reply was enthusiastic, and the founder built on a sample of one.
This is the honest gap in every guide including the good ones. Founders describe it directly:
"I need to talk to some doctors, chiefs of staff and hospital directors to see if the product is relevant at all, and to have a little team of 5 to 10 doctors that are willing to be presented as advisors on our pitch deck." – r/medtech
"I am trying mailing local gyms, beauty shops and clinics, but I do not know where else I can get customers. How does one get such clients at the beginning?" – r/freelancers
"I am trying to validate an idea I had recently. I want to scrape the internet, review sites, app stores, social media, to find what users think of your solution." – r/EntrepreneurRideAlong
That third one is a founder independently arriving at the reframe this page argues for, which is a reasonable sign it is the right instinct.
Three things work. First, go to where the complaints already are and contact the people who wrote them, because someone who publicly described a problem last week has demonstrated both that they have it and that they will talk about it. Second, use communities organised around the job rather than around a tool, since tool communities over-represent power users. Third, accept written exchanges. Ten thoughtful replies from real practitioners beat one scheduled call with a convenient friend, and they are an order of magnitude easier to get.
Our guides on finding business ideas on Reddit and using Reddit for idea validation cover the sourcing mechanics in detail.
Here is the reframe that changes the economics of discovery. Millions of people have already answered your questions in public, at length, unprompted, and in writing. They did it in reviews, in support threads and in forum posts. Nobody was pitching them. Nobody was watching their face.
The scale of what already exists is the argument. Across our corpus that is 39,000+ structured pain point records from review platforms, 9,000+ analysed software insights, 99,000+ negative app reviews, 40,000+ explicit feature requests, 29,000+ recorded product switches, and 2,300+ first-person accounts drawn from 180+ online communities. Every one of those is an answer to at least one question on this page.
| Question type | Answerable from written evidence? | What you still need a call for |
|---|---|---|
| Does this problem exist and recur | Yes, at scale | Nothing |
| What do people do instead | Yes, often in detail | The specific steps in your segment |
| What have they tried and abandoned | Yes, this is what reviews are | Why they picked that one first |
| Why did they switch | Yes, 29,000+ recorded instances | The internal politics of the decision |
| What do they complain about paying for | Yes, it is the largest complaint theme | What they would actually sign |
| Who approves the budget | Rarely | Almost everything |
| Would they try your rough version | No | All of it |
Three specific advantages. Volume, because you can read a thousand accounts in the time it takes to schedule two calls. Absence of performance, because nobody writing a review is trying to be encouraging. And retrospect, because the account is written after the fact by someone who now knows how the story ended, which is exactly the past-tense framing every good interview question is trying to induce.
The disadvantage is equally specific and worth stating plainly: you cannot ask a follow-up. That is not a small thing. It is the reason this is a first step and not a replacement.
A feature request is a complaint with a proposed solution attached. The complaint half is the reliable half, and the concentration across our feature gap records is striking enough to change where you point your discovery.
| Requested capability | Requests | Share rated high demand | What it means for your interview |
|---|---|---|---|
| Reporting | 2,899 | 69.5% | Ask what report they build by hand every week |
| User experience | 2,477 | 37.1% | High volume, lower intensity. Rarely a business on its own |
| Integration | 1,671 | 51.6% | Ask which two systems they retype data between |
| Analytics | 1,393 | 54.6% | Ask what question they cannot answer today |
| Data management | 660 | 55.3% | Ask where the spreadsheet lives |
| Financial management | 376 | 58.8% | Ask who reconciles, and how often it is wrong |
Reporting is both the most requested and the most intensely requested capability, at 69.5% high demand. User experience is the second most requested and among the least intense, at 37.1%. That contrast is the finding. People complain constantly about interfaces and rarely change software over them. They complain about reporting and then go build a spreadsheet, which is a workaround, which is a budget.
The same signal shows up in our validation data independently. Among the opportunity cards with enough exposure to be meaningful, the highest-scoring ones are dominated by reporting and analytics themes, with swipe rates between roughly 31% and 40%. Two independent sources pointing the same way is the convergence you want before committing discovery time.
Question sixteen is the highest-value question in most interviews, and it has already been answered at a scale no founder could reproduce. We hold 29,000+ competitive pairs extracted from public review text, each one recording that a reviewer described moving to or away from a named product.
| Incumbent | Mentions of switching away | Mentions of switching to | Ratio |
|---|---|---|---|
| HubSpot | 1,326 | 611 | 2.2 : 1 |
| QuickBooks | 1,137 | 579 | 2.0 : 1 |
| Salesforce | 948 | 401 | 2.4 : 1 |
| Asana | 907 | 446 | 2.0 : 1 |
| DocuSign | 728 | 281 | 2.6 : 1 |
| Zendesk | 714 | 332 | 2.2 : 1 |
| Trello | 594 | 307 | 1.9 : 1 |
Read the ratios rather than the counts. Every incumbent here shows roughly twice as many mentions of leaving as arriving, which mostly reflects that people write reviews when they are annoyed. What is useful is the variation: the spread from 1.9 to 2.6 marks which categories are in motion and which are settled. A market where customers are already moving is easier to enter than a static one, and your interview time is better spent with people mid-switch than with the contented.
More on this in why SaaS customers churn and turning review data into product ideas.
Question eleven is hard to ask well and easy to read. Across the 99,000+ negative app reviews we hold, 12.70% mention subscriptions, charges, refunds or paywalls. That is roughly double the 6.42% that mention crashes, bugs or freezing, and more than five times the 2.34% that describe something as confusing or hard to use. Money is the dominant complaint category by a wide margin.
The most useful of these state a pricing preference unprompted, which is precisely the answer question eleven is fishing for:
"Just another app that forces poor people to subscribe after 2 meds added. I would pay a one off price, no subscription for full access. Subscriptions target people who struggle with memory and organisation." – App Store review
"I used the free version for years. They decided to make it a paid service without giving any warning, now my inventory is locked up with no way to retrieve it. I would pay except they are 2 different apps." – App Store review
"This rating is for the free version. The ads are so intrusive as to render it useless. Not sure why I would pay for an app by people who thought this was a good idea." – App Store review
Three people volunteering the exact conditions under which they would hand over money, none of them asked. Only 2.53% of those negative reviews contain an explicit feature request, which is worth knowing: if you go looking only for people politely requesting features you will miss the other 97%, where the real information is.
"Barely functional anymore. Monthly view is stuck on August and July. Widget errors out regularly. Please add map view." – App Store review
For the pricing follow-through see SaaS pricing strategies and how to price a micro SaaS.
Four sources, each with a different bias, which is why you want more than one.
Review platforms. Business software buyers writing after months of use. Best for switching reasons and feature gaps. Biased toward categories with big marketing budgets. Our complaints database and guide to analysing review data cover the method.
App stores. Consumers, blunt, high volume, strongest signal on pricing and trust. Weakest transfer to business buyers. See how to analyse app store reviews and finding ideas in negative reviews.
Forums and communities. The richest narrative detail and the only place you reliably see the workaround described step by step. Our corpus spans 180+ communities, weighted toward the ones where practitioners actually gather: r/smallbusiness leads with 85 recorded pain points of which 75% are high impact, followed by r/teachers at 66, r/indiehackers at 45 and r/startups at 35 with 71% high impact.
Freelance marketplaces. The purest workaround signal, because every listing is somebody paying a human to do a thing manually. See the state of freelance demand for the current shape of it.
A note on reading forum threads specifically. The most valuable posts are rarely the ones asking a question. They are the ones describing a cost that has already been paid:
"I cannot leave work at work. Negative interactions stick with me and I struggle to not think about it. Quitting customer service has improved my mental health greatly." – r/CustomerService
"I kept getting rejected from transcription work until I realized this mistake. A lot of people quit gigs thinking it is skill when really it is rubric rules." – r/remotework
Neither is asking for a product. Both describe a recurring cost in detail, which is what question eight is trying to extract and what a politely answered interview usually fails to surface. The method behind reading these at volume is covered in finding business ideas on Reddit and Reddit research tooling.
Stop when new conversations stop producing new failure modes. In a narrow segment that is usually between eight and fifteen. If interview twelve tells you something you have heard four times, you are done with that segment. If interview twelve is still surprising you, your segment is too broad and you should narrow it rather than schedule more calls.
Sample composition matters more than sample size. Fifteen conversations with people who resemble you produce a confident wrong answer. Five conversations with people who have the problem worse than anyone you know produce a useful one. This is why question twenty-two, asking for an introduction to someone who has it worse, is worth more than several extra calls.
Score each problem on four axes after every conversation, while it is fresh. Frequency: does it recur on a schedule. Workaround: does one exist and what does it cost. Budget: is anyone already spending on it. Trigger: can they name the incident that would make them buy.
A problem that scores on all four is rare and should be pursued immediately. A problem that scores on frequency and workaround but not on budget is a real problem inside a business that will not pay, which is the most common and most expensive trap in this whole exercise. Our eight-stage validation framework and the validation checklist formalise this.
Suppose you are considering a reporting tool for small marketing agencies. Discovery in the order described above.
Read first. Reporting is the top requested capability in our feature gap data at 2,899 requests, 69.5% of them high demand. The forum accounts describe the workaround explicitly: twelve tabs, CSV exports, VLOOKUPs, four minutes a week per report. Validation cards in the reporting and analytics theme carry the highest swipe rates we record. Three independent sources agree before a single call.
Then interview, narrowly. You now do not need to ask whether reporting is painful. You ask which specific report, which twelve tabs, who asks for it, on what day, and what happened the last time it was late. Those are questions you could not have written before reading.
Then check the money. This is where the example turns. Marketing products in our revenue corpus have a median of $276 MRR across 213 revenue-reporting companies, with 28.6% clearing $1,000 MRR. That is one of the healthier categories we measure, so the willingness to pay is real, but the median tells you what a realistic outcome looks like rather than the best case. Our revenue benchmarks by category carry the full spread, and how to calculate market size shows how to turn it into a number.
| Minutes | What you are doing | Questions |
|---|---|---|
| 0 to 3 | Frame it as research, not a pitch. Say you are not selling anything | None |
| 3 to 10 | Get one specific recent story | 1 to 6 |
| 10 to 18 | Map the workaround in detail. This is the core | 7 to 10 |
| 18 to 24 | Follow the money and the approvals | 11 to 15 |
| 24 to 28 | Switching history, if the category has incumbents | 16 to 19 |
| 28 to 30 | Rank it, get an introduction, get a commitment | 20 to 24 |
Note what is absent. There is no slot for describing your idea. If it comes up, describe it in one sentence at minute 29, after every answer you care about has already been given.
Start with the interviews that already happened. BigIdeasDB holds 1M+ documented complaints, reviews and forum accounts, plus 40,000+ feature requests and 29,000+ recorded product switches, all searchable by problem rather than by product. Read first, then book calls that are worth the half hour.
Search the complaints →Every figure on this page was queried on August 31, 2026. Each source and what it cannot tell you.
| Source | Scale | Evidence type | Limitation |
|---|---|---|---|
| Review-platform pain points | 39,000+ records | Written complaints with severity | Reviewers skew unhappy. Volume tracks category popularity as much as category pain |
| Feature gap records | 40,000+ requests | Explicit unmet requests | Demand intensity is model-assigned from review text, not verified purchase intent |
| Competitive switch pairs | 29,000+ pairs | Stated migrations in review text | Mentions, not contracts. Churn pressure only, never market share |
| App store reviews | 99,000+ negative reviews | Unprompted consumer complaint language | Consumer behaviour transfers poorly to business buyers. Keyword shares are lower bounds, since people describe the same thing many ways |
| Forum pain points | 2,300+ records, 180+ communities | First-person narrative accounts | Community composition is not a representative sample of any market |
| Software insights | 9,000+ analysed insights | Structured review analysis | Sentiment labels are model-assigned and coverage is uneven across categories |
| Validation swipe data | 76,000+ opportunity cards | Human interest signal | Founders swiping, not buyers. Interest, not purchase intent, and only cards with real exposure are meaningful |
| Freelance demand | Job pain points by category | Paid manual work as a workaround proxy | Budget and rate fields are unpopulated, so no dollar figures are quoted from it |
| Revenue intelligence | 8,600+ startups, 3,700+ reporting revenue | Self-published revenue | Opt-in reporting, so absence is not zero and the sample skews toward willing disclosers |
All quotes are reproduced from public posts and reviews with usernames and post identifiers removed, attributed to the platform or community only. Light punctuation and spelling corrections were made for readability without changing meaning.
Written evidence cannot follow up. When a review says the reporting is useless, you cannot ask which report, and that follow-up is often the whole answer. Treat the corpus as the thing that tells you which follow-up to ask.
It cannot tell you who signs. Budget authority is almost never visible in public writing, and it is frequently the deciding factor in whether a real problem becomes a real purchase. That is call territory, permanently.
It over-represents the articulate and the annoyed. People who write reviews are not a random sample of users, and people who write long forum posts are not a random sample of professionals. Volume is a measure of how loudly a problem is discussed, which correlates with but is not identical to how badly it is felt.
And it is a snapshot. Everything here was queried on August 31, 2026. Complaint patterns move when incumbents ship, so re-run before relying on a specific share.
Next steps: validating a startup idea covers the full evidence ladder, idea validation collects the cluster, getting your first 100 users picks up where discovery ends, and finding problems worth solving runs the search against the same corpus quoted throughout this page.
The best ones ask about the past rather than the future. Walk me through the last time this happened. What did you do instead. How long did that take. What did it cost you. Who else was involved. What have you already tried and why did you stop. Every one of those has a verifiable answer. Questions like would you use this, or would you pay for this, ask someone to predict their own behaviour, and people are consistently generous and consistently wrong when they do that.
Fewer than most people think, if the questions are good, and more than most people run, because most people run zero. The practical rule is to keep going until new interviews stop producing new failure modes. That usually lands between eight and fifteen for a narrow segment. The bigger risk is not sample size, it is sample bias: fifteen conversations with people who resemble you produce a very confident wrong answer.
Anything hypothetical or flattering. Would you use this. Do you think this is a good idea. How much would you pay. Would you buy it if it did X. These invite politeness rather than information, and politeness is the default response to a founder describing something they clearly care about. Replace each one with a question about something that already happened, because the past is checkable and the future is a compliment.
It replaces the discovery half, not the relationship half. Public complaints, reviews and forum posts are unprompted accounts of problems written by people with no idea a founder was reading, which removes the politeness bias that damages live interviews. What they cannot do is let you ask a follow-up, and follow-ups are where interviews earn their keep. The efficient sequence is to read first and interview second, so you arrive already knowing the failure modes and can spend the call on the specifics.
This is the step where most discovery dies, and almost no guide addresses it. Founders in our forum corpus describe it plainly: one building medical software writes about needing to reach doctors and hospital directors just to find out whether the product is relevant at all. The workable answers are to go where the complaints already are and contact the people writing them, to use communities organised around the job rather than around the tool, and to accept that a written exchange with ten real practitioners beats a scheduled call with one convenient friend.
It is the rule that a good question is one your own mother could not lie about, popularised by Rob Fitzpatrick. You avoid pitching, avoid hypotheticals and ask only about specific past behaviour. It is the strongest single heuristic in the field. Its limit is practical rather than conceptual: it improves the questions but does nothing about getting people on the call in the first place, which is where most discovery actually fails.
Look for three things together. Frequency, meaning it recurs on a schedule rather than once. An existing workaround, because a workaround is a budget already being spent in time or money. And recency, because a problem people solved two years ago is not a market. In our review corpus the single most requested capability is reporting, with 2,899 recorded requests of which 69.5% are rated high demand, and it qualifies on all three counts across many categories at once.
Ask, but treat the answer as a description of a problem rather than a specification. Feature requests are valuable precisely because they are complaints with a proposed fix attached, and the complaint half is the reliable half. Across 40,000+ feature gap records we hold, requests concentrate heavily in reporting, user experience, integration and analytics. That concentration tells you where the pain is. It does not tell you that shipping the literal requested feature will solve it.
Related reading: business ideas that solve real problems, problem-solving app ideas, tools to find customer pain points, how to find startup ideas, daily frustrations that need an app and the idea validation tool.
BigIdeasDB Research. (2026). Customer Discovery Questions, and What the Answers Actually Look Like. BigIdeasDB. Retrieved from https://bigideasdb.com/customer-discovery-questions