The median market-research query is two words long. And the more precisely a founder describes what they are looking for, the less likely they are to find anything at all.
Founders search for market data the way they search Google. The median query in this study is two words long. That instinct is right, and almost everyone overrides it: the moment a founder tries to describe their actual niche precisely, their hit rate collapses.
We analyzed 34,000+ market-research queries run against the BigIdeasDB corpus between March and September 2026, of which 26,000+ carried free-text search terms. These are not survey answers about how founders say they research. They are the keystrokes, logged as they happened, from founders and the AI assistants they point at market data.
One-word queries return nothing 16% of the time. Four-to-six word queries return nothing 82% of the time. Precision makes you roughly five times more likely to find nothing. Search the broadest noun that still describes the problem, read what comes back, then narrow using the vocabulary in the results.
The relationship between how carefully a founder describes what they want and whether they find it is not weak, and it does not run the direction anyone expects. It inverts. Every extra word past the first makes the search worse, until the query becomes long enough to be a sentence and starts behaving like natural language again.
| Query length | Share of queries | Returned nothing |
|---|---|---|
| 1 word | 37% | 16% |
| 2 to 3 words | 40% | 54% |
| 4 to 6 words | 16% | 82% |
| 7 to 15 words | 5% | 48% |
| 16+ words | 1% | 80% |
A founder who types invoicing almost always gets something back. A founder who types invoicing software for independent contractors usually gets nothing. Same founder, same market, same corpus. The only thing that changed was how carefully they described it.
Every query in this study was run by a real person or their assistant against the BigIdeasDB corpus, through either the in-app research chat or the MCP server. Nobody was asked to perform a task. Nobody knew they were being measured. That makes it observational rather than experimental, which is both its strength and the source of its main limitation.
Those queries came from founders using the AI research chat, the MCP setup and the research server, so the sample spans both manual and assistant-driven research. The window runs March to September 2026 and covers 30 distinct research tools across complaint databases, funding data, Stripe company records and acquisition listings.
Across 26,000+ text queries the mean length is 3.7 words and the median is two. Roughly 37% are a single word. Another 40% are two or three. Only about 1% run past fifteen words.
This matters because it contradicts how research is usually taught. The advice is to be specific, define your segment, narrow your scope. Founders mostly ignore that when they sit down at a search box, and the data says their instinct outperforms the advice. The same broad-first instinct shows up in our guides to brainstorming business ideas and coming up with a business idea.
The dip at seven-to-fifteen words is the tell. Those queries are natural-language questions, and enough of their individual words land somewhere that they behave better than a tight four-word noun phrase. Then at sixteen words and beyond, failure jumps back to 80%.
So there are two workable modes and one trap. Broad nouns work. Full questions half-work. The precise middle, the four-to-six word description of exactly your niche, is where research goes to die.
A long phrase is matched as one literal string. For invoicing software for independent contractors to return anything, that exact sequence has to appear inside a stored record. Records are written by reviewers describing their own frustrations, so they never contain a founder's market definition verbatim.
One broad noun is different. invoicing appears in thousands of complaints, feature requests and category names, because it is a word the market itself uses. The founder is not being vague. They are speaking the corpus's language.
The most-searched terms in the dataset are unglamorous. They are the words a category would be named after, not the words a pitch deck would use.
| Term | Searches | Returned something |
|---|---|---|
| invoicing | 173 | 91% |
| real estate | 157 | 95% |
| automation | 152 | 98% |
| crm | 139 | 96% |
| scheduling | 137 | 96% |
| lead generation | 123 | 89% |
| job board | 93 | 70% |
| email finder | 73 | 56% |
automation, analytics, e-commerce, reviews, reporting, onboarding, monitoring. Each returned results between 98% and 100% of the time. They are all words that a software category, a review headline and a complaint would all plausibly contain.
email finder failed 44% of the time. job board failed 30%. enrichment 18%, automotive 21%. These describe the shape of a product someone intends to build, not a problem anyone complained about. Reviewers do not write "I need an email finder". They write that they cannot get contact data out of their CRM.
Two-word median, broad nouns, no operators, no segments. This is Google behaviour applied to a research corpus, and it is the correct instinct. It fails only when founders override it, usually right after deciding to get serious.
Reading the zero-result queries, almost all of them fall into three patterns. None is about missing data.
A founder searches for a category name that sounds official and is not. Retail Management System is a reasonable thing to call a category. The corpus calls it Retail Management Systems, plural, and holds no companies under it at all. The search was well-formed and still returned nothing.
The longest query in the dataset runs 316 words. It is an AI assistant pasting its entire working brief, hypothesis and approval boundary into a search box. Queries over fifteen words failed 80% of the time. A paragraph cannot match literally against anything.
This is the failure mode worth watching as more research runs through assistants. The model is not searching badly because it is stupid. It is searching badly because a search box and a prompt look identical to it.
Some of the highest-volume zero-result queries in the dataset differ from a real category name by one letter. That is not a founder mistake in any meaningful sense. It is a reminder that market research is still, mechanically, string matching. Our keyword research tool and the keyword generator guide exist largely to remove this class of miss.
The pattern underneath all three failures is the same. Founders search for the thing they intend to build. The evidence is filed under the thing users are annoyed about. Those are different vocabularies, and the gap between them is where research fails. It is the same mismatch we documented in customer support software limitations and email marketing limitations.
Adding words to a query feels like doing more careful work. It produces a more precise description of your own idea and a worse search. This is the single most expensive confusion in the dataset, and it explains why the failure rate peaks exactly where founders are trying hardest.
The same confusion shows up in how founders talk about validation. As one r/startups post put it: "Personal connection beats market research every single time." Not because research is useless, but because generic research produces generic conviction.
A serial founder writing in r/startups, in a post that drew over 1,300 upvotes, was blunt about where this ends: "it's incredibly easy to fall into a permanent research and planning phase without ever putting the rubber to the road."
And on scale: "I know tons of people who want to become entrepreneurs and start their own business and the research and planning phase is where the vast majority of them get stuck and never push past." His prescription is a time box: "time box your research to 10 hours for any opportunity."
If ten hours sounds short, note that the failure data says most of those hours were returning nothing anyway. For a structured version of the same discipline, see how to validate a startup idea and the validation checklist.
The query data shows the mechanism. Founders describe the consequence themselves, and they are unusually candid about it. A dev shop owner who has built MVPs for 25+ startups described the intake problem: "I asked what's the ONE thing this app needs to do? and he couldn't answer." The founder had a 47-page requirements document and no answer to the only question that matters.
The same post named the pattern that produces it: "The ones that failed ignored this advice and built in isolation for 6+ months." And the fix: "Not I'll do user research after I build it. Talk to them WHILE you're building."
An r/microsaas founder building for nightlife venues found out after shipping that he had researched an imaginary customer: "I built for the club I imagined. Nobody asked for multi-device sync." He had inferred the requirement from a plausible-sounding assumption and spent days on the edge cases.
Another, after eight products in three years, traced the split precisely: "The ones that failed I built because the idea seemed cool. Every single one." What replaced it was less romantic: "Boring problems genuinely make better businesses than interesting ones."
Research that stops at the wrong layer is its own failure mode. One r/SaaS founder researched a telehealth idea properly, talked to doctors, built the product, then hit a wall nobody had mentioned: "I found the barrier my optimistic, stupid mind never saw coming... Three months of work, gone."
And validation is easy to misread even when it arrives. A founder who got a surprise annual subscription could not tell what it meant: "I don't know whether they actually explored the app and saw value in it, or just paid for the first tool that came up on Google."
Benchmarks are misread the same way. As one r/microsaas founder pointed out about the growth posts everyone reads: "i keep reading these got my first 100 users in a week posts and every single one, when you dig in, turns out the person already had 8k twitter followers." More on that in getting customers and your first customer.
Small samples masquerade as signal too. One founder realised his roadmap was being set by a single user: "my most requested feature is from one guy who emails me every tuesday... 100% of that feedback was one person."
The successful accounts converge on a small number of moves, and none of them is more research. A serial founder writing in r/startups framed the search itself: "What we're looking for when doing research is companies in the space that are solving the same pain point we've identified." Not empty space. Company after company already solving it.
He is explicit that novelty is the wrong target: "it's pretty rare to have a truly new idea." Another put the same thing more bluntly after abandoning the search for originality: "demand has already been solved."
Difficulty selling is treated as data rather than as an obstacle to push through: "If you have to talk to one hundred people who are in your target market before anyone is even remotely interested..." The sentence is left unfinished in the original because the conclusion is obvious.
An investor-facing version of the same test came from a founder who reviewed hundreds of pitch decks: "Every deck should answer: what's the insight only you have? If I could've thought of your idea without domain expertise, it's not compelling enough." Generic research produces exactly the ideas anyone could have had.
The founder who crossed $50k on a Mac app ordered it plainly: "Validate Before You Build. Probably the most important point." And specified the test: "create a landing page, add a Stripe button, and try to sell the product before it exists."
He also drew the line most founders get wrong: "If you can build the product in two weeks, go ahead without validation. Otherwise, I recommend making sales before building." Research is a cost, and it should be proportional to the build.
One founder who reached 4 paying users in a day described a research process that reads almost exactly like the corpus method: "I started validating an idea across Reddit, forums, and Twitter, seeing if people would want AI meeting notes without sending conversations to the cloud." What worked: "talking about the problem, not the product." What did not: "overbuilding before validation."
Even unusual metrics count when they come from observed behaviour. The nightlife founder whose sync feature went unused found his real insight in the attendance data: "The 49% turn-up rate is the most valuable thing i noticed." A venue that knows its lists run at half can deliberately overbook.
And the market that stays unserved is worth noticing. One founder abandoned a working product after a friend criticised the code, and watched the gap persist: "To this day there is still no app like that locally, and that market is still not served." See lessons from failed business ideas and the road to first $1k MRR.
Finally, a signal about the tooling itself. An r/indiehackers post that drew heavy discussion argued the category has commoditised: "Every startup idea validator is AI now." Its author went the other direction and had real founders vote instead. The lesson generalises: an opinion generated from nothing is not evidence, whoever generates it.
The fix is not more effort. It is a different sequence. Start where the data is dense, let the corpus teach you its vocabulary, and only then narrow to your niche. Four moves, in order.
Type the single broadest noun that still describes the problem. Not your product, not your segment. invoicing, not invoicing for freelance designers. The failure data is unambiguous that this returns more, and the results are what tell you which narrower terms exist.
Reviewers describe frustrations, not solutions. Searching email finder looks for a product category; searching contact data or lead enrichment looks for the complaint. Our pain-point-first method and pain point tooling both start from this move, and the pain points database guide shows it in the product.
The first search is a probe for language. If scheduling returns records tagged appointment scheduling, medical scheduling and meeting, those are now your real query terms. This one habit removes most of the failures in this study, including the singular-versus-plural class.
An r/startups post makes the sharpest version of this case: "real validation is trying to kill your idea." The operational test it offers: "you ask what's the worst thing about how you solve this problem now? if they say nothing, your idea is dead." And a second screen: "you ask whose decision is this at your company? if they say not mine, you are talking to the wrong person."
The instruction that follows is the part most founders skip: "start looking for reasons your idea wont work." Our guide to niche viability and the multi-signal framework both formalize that instinct.
Once you can find records, the question becomes which records count. Three checks separate a real opportunity from an interesting one, and they map onto different corpora.
Demand is documented when strangers complained about it unprompted. That is what a review corpus is: 39,000+ scored pain points and 40,000+ feature gaps from Capterra, alongside 150,000+ G2 reviews and 136,000+ app-store reviews. See customer pain point analysis for the method and business pain points for the current picture.
A complaint with no budget behind it is a hobby. The Stripe Index holds 30,000+ companies verified live on Stripe across 83 categories, which answers whether money already moves in the niche. Freelance demand is a second read: see validating demand with Upwork jobs.
Incumbents existing is not the problem. Incumbents with no unfixable complaints is the problem. One r/microsaas founder described the exact filter: "I found mine by writing down everything that bugged me as a paying user of those tools, then crossing out anything they could fix in a week if they cared."
What survives that crossing-out is the wedge, and it is usually unglamorous. See the most-hated software of 2026 and sales software limitations for what that looks like at category scale. Our competitive landscape analysis, market gap analysis and the competitor analysis guide all run this test.
Most failed searches are aimed at the wrong corpus rather than at nothing. Pricing complaints live in reviews. Switching behaviour lives in competitive insights. Category saturation lives in company directories, not in review text.
Failure is not evenly spread. Searches aimed at company directories fail far more often than searches aimed at discussion data, because a directory is a closed list of names while a discussion corpus is open text. Knowing which you are querying changes how you should phrase it.
| What you are searching | Shape of the data | Returned nothing |
|---|---|---|
| App reviews | Closed list of apps | 75% |
| Companies on Stripe | Closed category taxonomy | 69% |
| Freelance opportunities | Role-titled categories | 70% |
| Software insights | Category-scoped text | 51% |
| Pain points | Open text | 48% |
| Subreddit fetches | Named source, no matching | 1% |
The bottom row is the control. Fetching a named subreddit fails 1% of the time because nothing is being matched, only retrieved. Everything above it fails in proportion to how much matching stands between the founder and the record. See finding app ideas from reviews and subreddit monitoring.
One caveat we would rather state than bury. Review coverage is not uniform across the 999 categories in the corpus. Some categories hold thousands of reviewed products and others hold almost none, which is an artifact of how the data was collected rather than a finding about those markets.
The practical consequence: a sparse result for a niche is ambiguous. It can mean the market is quiet, or it can mean we have not covered it yet. Treat a thin category as unmeasured rather than as evidence of a gap, and cross-check it against who is already selling there before drawing a conclusion.
| Your question | Best source | Scale |
|---|---|---|
| What annoys users in this category? | Complaint databases | 1M+ complaints |
| What features do they keep asking for? | Capterra feature gaps | 40,000+ records |
| Who already pays in this niche? | Stripe Index | 30,000+ companies |
| Where is capital flowing? | Funded DB | 17,000+ companies |
| What do products like this earn? | Revenue intelligence | 8,600+ startups |
| Is anyone selling one? | Acquisition listings | 650+ listings |
| What are people saying right now? | Reddit pain points | 2,300+ extracted |
1M+ complaints across Reddit, Capterra, G2 and the Apple App Store and Google Play, plus 1M+ embedded records for semantic retrieval. Depth per category is the number that matters for any single search, and it is uneven. See the state of SaaS pain points.
A review is written by someone who already paid, already used it, and was annoyed enough to write unprompted. A survey answer is a prediction about hypothetical behaviour. An r/SaaS founder who crossed $50k put it plainly: "I don't believe in surveys. Actual transactions are the strongest proof of demand and the ability to sell."
The same post carries the warning that makes research worth doing at all: "I've spent too much time building products that I ultimately couldn't sell, even though there was market demand." Demand existing and you being able to reach it are different findings. More in validating with real reviews and mining negative reviews.
Across 8,600+ revenue-verified startups the median monthly revenue is $0, and 4,900+ sit at exactly zero. That is the honest denominator. A research method is not good because it found something interesting; it is good if it moves you off that base rate. See the state of indie SaaS revenue and revenue intelligence.
A dev shop owner in r/SaaS described the shape of the failure: "Spent $120k and 5 months building. Launched. Got 31 signups." And the general case: "Most founders spend 6 months improving a product nobody wants."
An r/microsaas founder who shipped into analytics, a category everyone calls saturated, argued the opposite: "A crowded market isn't a bad sign. It means people pay for this stuff. The real question is whether the existing tools have annoying problems they can't easily fix."
That is the same test as check three, arrived at independently. See micro-SaaS competition, low-competition SaaS ideas and boring industries.
It cannot tell you whether these founders shipped, or whether better searching produced better companies. It observes queries, not outcomes. Anyone claiming a causal link from search behaviour to startup success is overreading this data, including us.
| Element | What we did | Limitation |
|---|---|---|
| Query dataset | 34,000+ logged research queries, March to September 2026 | Self-selected population: everyone here already chose a research tool, so they are not representative of all founders |
| Text queries | 26,000+ carrying a free-text term; length measured in whitespace tokens | Token counting treats hyphenated terms as one word, slightly understating length |
| Failure definition | A query returning zero records | A zero can mean bad phrasing, a genuinely absent niche, or thin coverage. We cannot always separate them |
| Corpus coverage | 1M+ complaints across five platforms, 999 review categories | Coverage is uneven by category. Some categories hold thousands of products and others hold none, so a thin result can reflect us, not the market |
| Human quotes | Public posts on Reddit (r/SaaS, r/startups, r/microsaas, r/indiehackers), anonymized | Reddit skews toward founders who post. Survivorship and self-promotion bias both apply |
| Revenue base rate | 8,600+ revenue-verified startups | Verified revenue skews toward founders willing to publish numbers, which likely overstates the median |
| Privacy | Aggregates only, rounded, no user-linked records | Raw logs are not publishable, so the underlying rows are not downloadable |
Three. The population is self-selected. The failure metric conflates phrasing problems with coverage problems. And observing queries tells you nothing about outcomes. The first and third are structural. The second we can partly separate, and intend to.
The two tables above are the study. Both are aggregate counts over a stated window with a stated failure definition, so any corpus with a query log can run the same measurement and compare. If your numbers disagree with ours, the interesting question is whether your corpus is organized around problems or around products.
Whether founders who adopt the corpus vocabulary after their first search go on to run more searches, and whether assistants can be taught to probe broad before narrowing. Both are measurable from the same log without touching outcomes we cannot see.
There is a second study sitting in the same data that we have deliberately not run yet: whether the categories founders search for most are the categories that actually support revenue. The two are not obviously the same. The most-searched terms in this study are broad horizontal categories, while the products that clear meaningful revenue in our valuation data skew narrow and vertical. If that divergence holds, the popular search is a poor guide to the good market, which would make vertical niche research and profitable niche selection more valuable than category browsing, not less. We will publish that either way.
The fastest way to skip the two-word learning curve is to let the corpus pick the category for you. Paste your product URL and the free teardown matches you to a category and returns what reviewers in it actually complain about.
34,000+ queries run against the BigIdeasDB corpus between March and September 2026, of which 26,000+ contained free-text search terms. Every query came from a real founder or their AI assistant researching a market, not from a survey or a lab task.
The precision paradox. A one-word query returns nothing 16% of the time. A four-to-six word query returns nothing 82% of the time. Describing what you want more precisely makes you five times more likely to find nothing at all.
Two words. The median across 26,000+ text queries is two words and the mean is 3.7. Roughly 77% of all queries are three words or fewer, so founders search market data the way they search Google, not the way they would brief an analyst.
Because a longer phrase is matched as one literal string. A six-word description of your exact niche has to appear verbatim in a stored record to match anything, and it almost never does. Short broad nouns match thousands of records; long precise phrases match none.
Search the broadest noun that still describes the problem, then narrow using what comes back. Start with invoicing, scheduling or recruiting rather than a full sentence describing your product. Broad terms in our data return results 86% to 98% of the time.
One-word category nouns. In our data automation returned results 98% of the time, analytics and e-commerce 100%, crm and scheduling 96%, and invoicing 91%. All of them are terms a category would plausibly be named after.
Compound product descriptions. Email finder returned nothing 44% of the time and job board 30%, because both describe a product shape rather than a category. The corpus is organized by the problem, not by the product you have in mind.
Usually not. Most zero-result searches in our data were for niches the corpus does hold, phrased in a way that could not match. That is a search-behaviour problem, and it is why the fix is a change in how you query rather than more data.
No, they fail differently. A meaningful share of the longest queries in the dataset were AI assistants pasting an entire brief into a search box, some over 300 words. Those failed 80% of the time, because a 300-word paragraph matches nothing literally.
Less than most founders think. One serial founder in r/startups recommends time-boxing research to ten hours per opportunity, on the grounds that the research and planning phase is where most would-be founders get permanently stuck.
No. A crowded market proves people pay for the category. The question is whether incumbents have problems they cannot cheaply fix, which is exactly what documented complaints and feature gaps reveal.
Validation looks for reasons to proceed and finds them easily. Invalidation looks for the reason the idea dies. As one r/startups post put it, real validation is trying to kill your idea, and the fastest kill shot is asking what the worst thing about the current solution is.
A review is written by someone who already paid, already used the product, and is annoyed enough to write it down unprompted. A survey answer is a prediction about hypothetical behaviour. One founder in r/SaaS put it bluntly: actual transactions are the strongest proof of demand.
That the median outcome is zero. Across 8,600+ revenue-verified startups the median monthly revenue is $0, with 4,900+ sitting at exactly zero. Any market-research method has to be judged against that base rate.
It observes queries against one corpus, from a self-selected population of founders who already chose a research tool. Corpus coverage is also uneven by category, so a thin result for a niche can reflect our coverage rather than the market. Both caveats are in the methodology table.
The aggregate query-length and zero-rate tables are reproducible from the figures in this article. We publish rounded aggregates and dated snapshots rather than raw logs, because the raw logs contain user-linked records that should never be published.
Run one broad category noun, read what comes back, then use the vocabulary in those results as your next query. If you want the same thing done for the category you already compete in, the free site teardown does it from your URL.
Related reading: how to find startup ideas, SaaS ideas backed by pain points, customer discovery questions, market sizing, customer review analysis, competitor research tools, Reddit research tools, SaaS research tools, AI for market research, market research tools, the idea validation playbook, validating a side hustle, finding problems to solve, idea validation tools, micro-SaaS ideas, AI product validation, finding a profitable niche, mining Capterra reviews, turning G2 reviews into ideas, app store review analysis, Reddit market research, the complaint analysis platform, researching market size, the idea validation tool, SwiperDB, and indie hacker validation. Start with the pain points database or Discover.
BigIdeasDB Research. (2026). How Founders Actually Research Markets: 34,000+ Queries Analyzed. BigIdeasDB. Retrieved from https://bigideasdb.com/how-founders-research-markets