A nine-criterion rubric where every data point comes from a live source, a worked example with three real idea categories, and the go-deep sprint that settles a tie.
To choose between startup ideas, score each one on the same nine criteria, weight the ones that predict revenue and reach, pick the top score, and then commit hard. Seven criteria come from market data: complaint volume, complaint severity, documented market gap, saturation, revenue proof, capital flow and buildability. Two come from you: unfair advantage and distribution access. When we ran three plausible SaaS ideas through it with live data, complaint volume differed by more than 30x and builder density ranged from 0.4% to 34.7%. The ideas were not equal. They only looked equal from inside the founder’s head.
This page is for founders who already have ideas. Most are experienced: people who have shipped before, hold three or four candidates, and have exactly one build slot for the next six months. If you are still looking for a first idea, start with what to do when you have no ideas or how to find startup ideas. If you are choosing what kind of business to start at all, how to decide what business to start covers that broader question. This one ranks SaaS ideas against each other.
Every number here comes from BigIdeasDB’s warehouse, queried read-only on September 25, 2026: a complaint corpus of 1M+ records, 30,000+ companies from Stripe’s public directory, 8,600+ revenue-verified startups and 17,000+ funded companies. Where a number is thin, we say so.
The rest of this page explains each criterion, shows the exact data behind it, walks through a full worked example, and covers what to do once the decision is made. That last part matters most. A good choice executed shallowly loses to an average choice executed deeply.
This rubric assumes you can build or get something built, you have some runway, and your constraint is focus rather than ideas. That describes most repeat founders and a lot of senior engineers leaving jobs. One r/startups post from an engineer after a failed first company put the constraint in one line:
“I don’t have a strong B2B network yet, and I’m not naturally great at sales.” – r/startups
That sentence is a rubric input, not a confession. It should lower the distribution score of any idea that depends on enterprise sales. Another founder, deciding between two marketplace ideas, named the one-slot constraint directly:
“I can build one proper MVP and want to choose carefully rather than spend months building something with weak economics.” – r/startups
If you do not yet have candidates, this page will feel premature. Use how to come up with a business idea or browse Discover first, then come back with three.
The problem is not a shortage of ideas. It is that unexamined ideas cannot be compared. One r/indiehackers founder counted the backlog honestly:
“I just counted. I have 17 ideas in my notes. Some are three-year old. Most are literally 2-3 sentences.” – r/indiehackers
And then named exactly why choosing felt impossible:
“You can’t compare a half-formed thought about a fitness app to a half-formed thought about a newsletter business. They’re both just vague possibilities. There’s nothing to compare.” – r/indiehackers
Experienced founders add a second problem. They know enough to see risks in every option, so every option looks flawed. One r/startups founder who runs several threads at once described the loop:
“The hard part isn’t generating ideas; it’s deciding what deserves attention today and not reopening the whole strategy every time I sit down.” – r/startups
A rubric fixes both problems. It turns vague ideas into comparable rows, and it gives you a recorded reason for the choice so you stop relitigating it. The SaaS idea validation tool collects the raw evidence for each row. For the single-idea version of this discipline, see how to stop delusional thinking when validating.
In June 2026 Y Combinator published Pick One Idea and Go Deep, a Startup School talk by YC General Partner Jon Xu. It had 162,000+ views by late September. Its opening diagnosis is the reason this page exists:
“I often meet founders who have lots of ideas about what to work on and can’t decide between them.” – Jon Xu, Y Combinator
His prescription has two halves. First, stop overthinking: there is no perfect idea to find in the abstract. Second, once you choose, commit completely:
“You should burn the other boats. That is, you should explicitly foreclose your other startup idea options. Stop working on them.” – Jon Xu, Y Combinator
The rubric on this page sits between those halves. It is a fast, evidence-based way to make the pick so you can move to the part YC actually cares about: depth. His test for depth is blunt: could you run your customer’s business? Could you teach a class on the problem? That is the standard your chosen idea gets held to after this page.
The tempting alternative to choosing is testing everything at once. The YC talk explains why that fails:
“If you don’t actually go deep on an idea, but instead juggle it with several others, you won’t get good signal about whether what you’re doing actually works.” – Jon Xu, Y Combinator
Shallow tests generate false negatives. A landing page that got 40 visits did not test demand. It tested nothing. Founders then drop a good idea because it “did not work” after a weekend. One r/SaaS founder who did go all in on two ideas at once got lucky, and said so:
“I quit my job without knowing which one would work. Went all in on both.” – r/SaaS
The one that worked was a simple service with a risk-free offer. The AI product “died quietly.” Running both only worked because one idea needed almost no building. Most SaaS ideas are not like that. A short disconfirmation test per idea is fine. Parallel building is not.
A build slot is the scarcest thing a small founder owns. It is the next three to six months of focused engineering and selling. Spend it on the wrong idea and the cost is not just the idea. It is every compounding week the right idea did not get. A serial founder with nine startups behind them wrote a 1,400+ upvote r/startups guide on picking ideas and ended it here:
“You probably have a bunch of other ideas you are excited about the potential of and are interested in exploring... But unfortunately, your limited resource is always time.” – r/startups
The base rates make the slot even more precious. Of 8,600+ startups tracked in TrustMRR as of September 2026, only 43.5% earn any revenue at all. Of those that do, the median is $145 in monthly recurring revenue and 6.1% reach $10,000. About a quarter (24.1%) are listed for sale. For the wider picture see the state of indie SaaS revenue and TrustMRR revenue benchmarks. A bad pick is the default outcome, not the exception.
The rubric will not find a perfect idea, and you should not want it to. Both YC talks on this topic make the same point. In How to Get Startup Ideas, YC partner Jared Friedman notes that Google was roughly the 20th search engine and Facebook roughly the 20th social network. What made them work was a good-enough idea plus execution.
Friedman also names the opposite mistake: jumping into the first idea without spending even a couple of weeks deciding. His point is that you will spend years on a successful startup, so a few weeks of choosing is cheap. The rubric is sized for that window. It should take days, not months, to fill in.
The practical rule: aim for a clearly better choice, not the best possible one. The goal is a good starting point that can morph, which is also why startup pivot examples so often start from a reasonable first idea rather than a brilliant one.
Each criterion is scored 1 to 5. Seven map to a specific data signal you can pull today. Two are about you and have no database behind them, which is exactly why they get heavy weights: they are where most outcomes are decided.
| # | Criterion | The question | Data signal | Weight |
|---|---|---|---|---|
| 1 | Complaint volume | How many people describe this problem unprompted? | Cross-corpus mentions in the 1M+ complaint corpus | x2 |
| 2 | Complaint severity | How much does it hurt when it happens? | Capterra pain severity, Reddit impact rating | x1 |
| 3 | Documented market gap | Do existing tools fail at this? | Capterra market-gap and competitive-gap scores | x2 |
| 4 | Saturation | How many builders like you are already here? | Stripe Index micro-SaaS density | x2 |
| 5 | Revenue proof | Do products in this category actually earn? | TrustMRR median MRR and share above $1,000 | x3 |
| 6 | Capital flow | Is money moving into this space? | Funded DB momentum score | x1 |
| 7 | Buildability | Can a small team ship it? | Stripe Index buildability score, Capterra feasibility | x1 |
| 8 | Unfair advantage | What do you know that others do not? | None. Self-assessed | x3 |
| 9 | Distribution access | Can you reach 20 buyers without ads? | None. Self-assessed | x3 |
If you know the SaaS opportunity score or the seven-signal validation method, criteria 1 to 7 will look familiar. The difference is purpose. Those tools judge whether one idea is worth pursuing. This rubric ranks several ideas against each other for a single slot, and it gives founder fit explicit weight instead of treating it as a footnote.
Score data criteria against fixed thresholds, not against each other. Relative scoring hides a weak field: if all three ideas are poor, a relative score still awards someone a 5. Fixed thresholds also make your scores comparable over time, so the idea you park today can be rescored in six months.
| Criterion | 1 | 3 | 5 |
|---|---|---|---|
| Complaint volume (cross-corpus mentions) | Under 200 | 500 to 2,000 | Over 3,000 |
| Complaint severity | Mostly low impact | Mixed | Most mentions high impact |
| Market gap (Capterra, 0 to 10) | Under 4.5 | 5.5 to 6.5 | Over 7.5 |
| Saturation (micro-SaaS share) | Over 25% | 8% to 15% | Under 2% |
| Revenue proof (median paying MRR) | Under $100 | $150 to $300 | Over $600 with a solid sample |
| Capital flow (momentum, 1 to 10) | Under 4.5 | 5.0 to 5.9 | Over 6.5 |
| Buildability | Regulated or hardware | Standard SaaS with integrations | Single-feature, no dependencies |
| Unfair advantage | None | Relevant adjacent experience | Years inside this buyer’s job |
| Distribution access | Cannot name 5 buyers | Can name 20 with effort | Existing audience or customers |
Complaint volume answers one question: how many strangers describe this problem without being asked? It is the cheapest demand signal there is, because the people wrote it for other reasons. BigIdeasDB’s corpus holds 1M+ records, including 39,000+ structured Capterra pain points, and 99,000+ negative app store reviews as of September 2026.
To score it, count mentions across several independent sources, not one. In our worked example we counted Capterra review cons, Reddit pain points, G2 insights and Capterra feature gaps. The totals were 120+, 800+ and 4,100+ for the three ideas. That spread is typical: volume is the criterion most likely to separate ideas that feel equal. Start in the pain points database; the pain points guide covers the filters.
One caution. Volume rewards big, old categories. A new problem can be real and still have few mentions. That is why volume is weighted x2, not x3, and why severity sits beside it.
Severity asks how much the problem costs the person when it happens. A mild annoyance mentioned often is worth less than a costly failure mentioned less. Capterra pain points carry a severity score, and Reddit pain points carry an impact rating.
In practice severity separates less than you would expect. Across our three worked-example ideas, average Capterra severity ranged only from 3.82 to 3.89 on a 5-point scale. The share of pain points at severity 4 or above was more telling: 55.8%, 65.1% and 62.4%. On Reddit, 81.5% of the bookkeeping-related pain points were rated high impact, against 73.6% for trades. Severity is a tie-breaker inside demand, which is why it carries x1.
The quickest way to feel severity is to read the quotes. A Capterra reviewer on an accounting add-on:
“Big drawback: it doesn’t connect to QuickBooks. Without a native sync, we’re stuck doing duplicate data entry and manual reconciliations.” – Capterra review
That is a severity-4 complaint in plain language: recurring manual work caused by a missing integration. For more on reading reviews this way, see customer review analysis.
Volume and severity tell you the problem is real. The gap criterion tells you existing tools are failing at it. BigIdeasDB’s Capterra layer scores this two ways: a market-gap score on systemic category pain points, and a competitive-gap score on 3,100+ pre-analyzed SaaS opportunities. Only 380+ of those opportunities score 7 or higher overall, about 12%, so a high score here is genuinely selective.
In the worked example, average category market-gap scores were 5.35, 6.23 and 6.97. Competitive-gap scores followed the same order: 4.98, 5.67 and 6.22. The Capterra analysis guide shows where to find both, and the most underserved software markets ranks categories on the same idea.
Read the top gaps, not just the average, the way business pain points in 2026 does. For the AI writing group, the highest-gap category pain points were about pricing: “Inflexibility in Pricing Structures” and “High Subscription Costs Lacking Value for Money.” For the trades group, one was “Lack of Integration with Existing Systems for QuickBooks Syncing.” A pricing gap invites a price war. An integration gap invites a product.
Most founders measure competition by counting companies. That is the wrong count. What matters to a small founder is how many other small software products already serve the same buyer, because those are the products you fight for attention, reviews and search results.
The Stripe Index classifies 30,000+ companies that take payment through Stripe, and flags 2,000+ of them as micro-SaaS. Across the whole index, micro-SaaS is 6.6% of companies. Inside categories it varies wildly. Of 66 categories with at least 50 companies, 10 run under 2% micro-SaaS and 13 run over 15%.
| Category | Companies | Micro-SaaS | Density | Saturation score |
|---|---|---|---|---|
| AI Tools & Apps | 950+ | 330+ | 34.7% | 1 (builder-crowded) |
| Accounting & Bookkeeping | 280+ | 30+ | 12.1% | 3 |
| Scheduling & Booking | 2,000+ | 100+ | 5.0% | 4 |
| Home Services & Trades | 960+ | Under 5 | 0.4% | 5 (operators, not builders) |
| Whole index | 30,000+ | 2,000+ | 6.6% | Baseline |
AI Tools and Home Services hold almost the same number of companies. One is a third builders. The other is almost entirely operators with no small software serving them. For the full map, read SaaS market saturation, micro SaaS ideas from the Stripe Index and the state of micro SaaS competition, or query it directly from the Stripe Index database.
Competition is not the enemy here. A category with real paying companies and few builders is the best of both: proof of spend without a crowd.
Revenue proof is the heaviest data criterion because it answers the question every other signal only implies: do products like this actually earn? TrustMRR tracks 8,600+ startups with revenue read from connected payment providers, not founder claims. For each idea, pull the median monthly recurring revenue of paying startups and the share that clear $1,000.
Index-wide, paying startups median $145 MRR and 23.6% clear $1,000. Use those as the bar. In our worked example, AI writing tools medianed $190 with 24.1% above $1,000, which is roughly the index. Trades software medianed $500 with 36.8% above $1,000, but on fewer than 20 paying startups. Bookkeeping tools medianed $113 with only 8.1% above $1,000. The TrustMRR guide shows how to filter to your idea.
Sample size matters more here than anywhere. A median from under 20 startups is a hint. A median from nearly 200 is evidence. We cap the score at 4 when the paying sample is under about 30, and you should too. For per-niche revenue context see the most profitable SaaS niches. For category ceilings in more detail see what micro SaaS actually charges.
Capital flow asks whether investors are backing companies in the space. BigIdeasDB’s funded company database scores 17,000+ companies on momentum from 1 to 10. The index average is 4.91.
It is the criterion founders overweight most. In our worked example, momentum for companies matching the three ideas was 5.48, 5.32 and 5.45. All slightly above average, none separable. That is common: venture money follows broad themes, and most bootstrappable SaaS ideas sit in themes that are neither hot nor cold. Capital flow matters when it is extreme, such as a category with no funded activity at all, or a surge you can ride. Otherwise it is a tie-breaker, weighted x1.
If you plan to raise, read it differently: see how to find startup ideas that get funded and browse the funded startups database. The Funded DB tools guide covers querying it from an AI assistant.
Buildability asks whether a small team can ship a credible first version. The Stripe Index scores every company’s product from 1 to 10 on how buildable it is for a small team, and Capterra opportunities carry an implementation-feasibility score.
In 2026 this criterion barely separates SaaS ideas. The average buildability of micro-SaaS products already in our three categories was 6.34, 6.50 and 6.41. Across all micro-SaaS in the index it is 6.52. Modern tooling has flattened the difference. Feasibility showed a little more spread: 4.80, 5.09 and 4.23, with bookkeeping lowest because accounting products live or die on integrations and correctness.
The trap is treating easy-to-build as a reason to choose. One r/SaaS founder described where that leads:
“AI makes it so easy to build now that I keep shipping things nobody wanted.” – r/SaaS
Buildability gets x1. If you are non-technical, weight it higher and read what to build as a solo developer for the constraints that actually bite.
No database can score this one. Unfair advantage is what you know, or who you know, that a competent stranger does not. YC’s Jared Friedman calls this founder/market fit and says roughly half of YC’s most successful companies trace back to one recipe: start with what the team is especially good at and pick ideas where that gives an unfair advantage.
Score it honestly and narrowly. “I am a good engineer” is not an advantage for any specific idea. “I spent four years running dispatch for an HVAC company” is a 5 for a trades idea and a 1 for everything else. Interest counts a little. One founder on r/Entrepreneur described why:
“Personally I choose the one that plays to my strengths. Because once you start going into a rabbit hole, that’s what keeps you going on the boring days.” – r/Entrepreneur
The YC talk adds a warning for repeat founders: do not use this criterion to disqualify yourself. Xu says second-time founders often “weaponize” founder-market fit, convincing themselves they need a decade of domain experience. Deep customer work can build real expertise fast. So score today’s advantage, but remember it can change. For idea-by-background examples see validating an idea in an industry you do not know.
Distribution access is the criterion that most often decides a small company’s fate and the one founders most often skip. The test is concrete: can you name 20 specific buyers you could contact tomorrow, without paid ads? Not “accountants,” but 20 named firms.
A serial founder’s guide on r/startups frames the same test from the other side:
“If you have to talk to one hundred people who are in your target market before anyone is even remotely interested... it’s a good sign you are not in an opportunity you should chase.” – r/startups
Existing customers, an audience, a community you belong to or a partner who sells to your buyer all push this score to 4 or 5. If you are starting cold, it is a 1 or 2, and that is fine as long as you know it. For how distribution plays out after launch, read how to get your first customer and the growth levers founders never pulled.
Equal weights assume every criterion predicts success equally. They do not. Three criteria are weighted x3 because each one can end a company by itself: no revenue in the category, no edge, or no way to reach buyers. Volume, gap and saturation get x2 because they shape how hard the fight will be. Severity, capital flow and buildability get x1 because in practice they rarely separate SaaS ideas.
The weighting is also a defence against the most common founder error. CB Insights’ analysis of 431 VC-backed shutdowns since 2023 found 43% cited poor product-market fit. Running out of capital topped the list at 70%, but CB Insights calls it the final cause, not the root. Weighting revenue proof and distribution heavily pushes you toward ideas where fit is already partly visible.
Change the weights if your situation demands it. A founder with 18 months of runway can down-weight buildability. A founder who must reach revenue in 90 days should weight speed to first dollar inside revenue proof. One r/startups founder described exactly that shift:
“When i had runway, build speed mattered less. when i was bootstrapping, build speed was everything because i couldnt afford to spend 3 months validating before seeing a dollar.” – r/startups
Here is the rubric run end to end on three idea categories a repeat SaaS founder might plausibly hold at once. We chose them because they are common, they sit in categories where every data source has coverage, and they pull in different directions.
Each idea maps to Capterra categories (for example, Artificial Intelligence, Content Marketing and AI Writing Assistant for A; Field Service Management, HVAC, Landscape and similar for B; Accounting, Bookkeeper and Billing and Invoicing for C), to one Stripe Index category, and to keyword matches in TrustMRR startup descriptions and funded company one-liners. The exact definitions are in the methodology. If you want the reporting-and-analytics version of a multi-source example, multi-signal validation has one.
| Signal | A: AI writing | B: Trades field service | C: Bookkeeping |
|---|---|---|---|
| Capterra review cons mentioning the problem | 70+ | 530+ | 2,700+ |
| G2 insights mentioning it | 15+ | 150+ | 530+ |
| Capterra feature gaps mentioning it | 20+ | 60+ | 690+ |
| Reddit pain points mentioning it | 10+ | 50+ | 100+ |
| Cross-corpus total | 120+ | 800+ | 4,100+ |
| Capterra pain severity (1 to 5) | 3.82 | 3.86 | 3.89 |
| Share of pain points at severity 4+ | 55.8% | 65.1% | 62.4% |
| Category market-gap score (0 to 10) | 5.35 | 6.23 | 6.97 |
| SaaS opportunity competitive gap | 4.98 | 5.67 | 6.22 |
| Stripe Index micro-SaaS density | 34.7% | 0.4% | 12.1% |
| Median MRR, paying startups | $190 | $500 (small sample) | $113 |
| Share of paying startups above $1,000 MRR | 24.1% | 36.8% | 8.1% |
| Paying startups in sample | 190+ | Under 20 | 30+ |
| Funded momentum (index 4.91) | 5.48 | 5.32 | 5.45 |
| Micro-SaaS buildability (1 to 10) | 6.34 | 6.50 | 6.41 |
| Implementation feasibility | 4.80 | 5.09 | 4.23 |
Three things jump out before any scoring. Complaint volume differs by more than 30x between A and C. Density differs by almost 90x between A and B. And the ideas that are loudest in complaints are not the ones that earn best: bookkeeping has the most documented pain and the weakest revenue medians. That tension is exactly why you need several signals instead of one.
The complaint text shows the difference in kind, too. A Capterra reviewer of an AI writing tool:
“Some of the AI-generated text/keywords still felt generic and needed editing to fit our brand voice.” – Capterra review
A reviewer of a field service product:
“I cannot export the driver files in CSV format. I would also like to have the option of dispatching another route before the previous one is completed.” – Capterra review
The first is a quality complaint about a commodity. The second is a specific workflow gap a small product could own.
Applying the thresholds above gives the data half of the rubric. Founder-fit scores are for a hypothetical founder: a former product manager at a marketing SaaS company, with a network of marketing leads and no trades or accounting background. That founder is common among our users, which is why we chose them.
| Criterion (weight) | A: AI writing | B: Trades | C: Bookkeeping |
|---|---|---|---|
| Complaint volume (x2) | 1 → 2 | 3 → 6 | 5 → 10 |
| Complaint severity (x1) | 3 → 3 | 4 → 4 | 4 → 4 |
| Market gap (x2) | 2 → 4 | 3 → 6 | 4 → 8 |
| Saturation (x2) | 1 → 2 | 5 → 10 | 3 → 6 |
| Revenue proof (x3) | 3 → 9 | 4 → 12 (capped) | 2 → 6 |
| Capital flow (x1) | 3 → 3 | 3 → 3 | 3 → 3 |
| Buildability (x1) | 4 → 4 | 4 → 4 | 3 → 3 |
| Data subtotal (of 60) | 27 | 45 | 40 |
| Unfair advantage (x3) | 4 → 12 | 1 → 3 | 2 → 6 |
| Distribution access (x3) | 5 → 15 | 1 → 3 | 2 → 6 |
| Total (of 90) | 54 | 51 | 52 |
On market data alone, trades software wins clearly: 45 against 40 and 27. Add this founder’s fit and the three ideas finish within three points of each other, with the AI writing tool nominally first. That is not a failure of the rubric. It is the most useful thing it can tell you.
The market says B. This founder’s network says A. Run the same rubric for a different founder, one who spent years in operations at an HVAC company, and B scores 5 on both founder criteria. Its total jumps to 75 while A drops to around 39. Same market, same data, opposite answer.
This is why a rubric without founder fit is incomplete, and why a rubric that is only founder fit is dangerous. The data tells you how hard each market will be for anyone. Founder fit tells you how hard it will be for you. A founder on r/Entrepreneur got at the same split:
“What do you like most and what would the market support.” – r/Entrepreneur
The honest reading for our hypothetical founder is that A is winning on their network, not on the market. Its data subtotal of 27 is the weakest of the three. If that network does not convert into paying customers quickly, nothing else in A’s numbers will rescue it.
Before you trust a total, check how fragile it is. Change one input at a time and see whether the ranking moves.
| Change | New totals | Winner |
|---|---|---|
| Baseline | 54 / 51 / 52 | A, narrowly |
| Drop founder fit entirely | 27 / 45 / 40 | B |
| Equal weights on all nine criteria | 26 / 28 / 28 | B and C tie |
| Founder gets distribution to 3 for trades (e.g. a partner who sells to contractors) | 54 / 57 / 52 | B |
| Revenue proof for B scored 3 instead of 4 | 54 / 48 / 52 | A |
| Founder is an ex-HVAC operator | 39 / 75 / 46 | B, clearly |
The pattern: B wins whenever the founder can move either founder-fit score up by even two points. A wins only when the founder’s existing network is the deciding factor. That tells our hypothetical founder exactly what to test.
When two ideas finish within about 10% of each other, the rubric has done its job by removing the weak options. Now break the tie on fixability: pick the idea whose weakest criterion you can change fastest.
Applied here, the tie-break favours B, provided the founder commits to closing the founder-fit gap. That matches the YC talk’s advice to stop disqualifying yourself for lack of domain experience. One r/Entrepreneur commenter described the moment you are in:
“One pattern I see here is that the problem is actually not having many ideas, but that none of them has crossed a decision threshold yet.” – r/Entrepreneur
A near-tie means both ideas crossed the threshold. The rest is commitment.
The rubric picks. The sprint confirms. Two weeks, one idea, no building beyond what a test needs. This is the practical form of YC’s “go deep”:
YC’s bar for when the sprint has gone deep enough: could you run your customer’s business? In the talk’s example of voice agents for cleaning companies, the question is not whether you talked to 20 owners but whether you know their daily crises and what they would pay to never lose another call. For the longer validation path, see the 8-stage validation framework and how to validate a startup idea.
Kill criteria are the conditions under which you will drop the chosen idea, written before the sprint so excitement cannot move them. One r/startups founder shared a version with a hard cut-off:
“Score five things from 1 to 5: pain, existing spend, reach, build speed, monetization. I add them up out of 25. If total is under 14 I kill unless I have new evidence.” – r/startups
Good kill criteria are observable and dated. Examples that work:
Another founder in the same thread offered a shorter version:
“If i cant describe the specific person who would buy this AND explain why they havent solved this problem already with existing tools, its dead.” – r/startups
If the idea dies, that is a success of the process, not a failure of the founder. See lessons from failed business ideas for what the pattern looks like when founders skip this step.
Every founder knows friends are too kind. Few act on it. One r/startups founder described the gap between what people say and what they do:
“Ive had at least two ideas where 5/5 people in conversations said theyd buy it and exactly zero did when it shipped. the gap between stated and revealed preference is enormous.” – r/startups
Another put it more bluntly:
“The 5 conversations one i’d be careful with, friends of friends are basically a rigged jury.” – r/startups
This is the practical case for the data half of the rubric. Complaint volume, revenue medians and saturation are stranger signals at scale: they were produced by people with no interest in your idea. Use them to choose, then use real conversations to confirm. For how to read public signals well, see validating a SaaS idea with real reviews and validating demand with Upwork jobs.
The biggest risk of any scoring system is that it dresses up the decision you already made. A founder on r/startups admitted it directly:
“The scoring system would just launder my own bias tbh. If i’m excited about an idea i’ll quietly bump pain to a 5 and reach to a 4 and suddenly it clears 14.” – r/startups
And a co-founder who had killed three product lines agreed:
“When you’re emotionally attached, the scoring system becomes a rationalizer.” – r/startups
Four rules keep it honest:
The startup idea validation checklist is a useful companion for the evidence you should collect under each score.
In our worked example, the unglamorous trades idea had the best market data. That is not a coincidence. Categories full of operators and empty of builders stay that way because builders find them dull. YC’s Friedman lists “ideas that are in a boring space” as one of four filters that make founders reject good ideas unconsciously, citing payroll software.
The r/SaaS founder whose AI product flopped while a simple service took off drew the same conclusion:
“The idea that felt too simple, too obvious, too unsexy was the one that actually had demand behind it.” – r/SaaS
A commenter summed it up:
“‘Impressive’ ideas and ‘valuable’ ideas are often completely different things.” – r/SaaS
If you want a list of these spaces, start with niche SaaS opportunities by industry, boring business ideas and small business software pain points. The same logic drives vertical AI SaaS ideas: dull verticals with real budgets.
AI ideas are not bad. They start with a structural handicap on the rubric. AI Tools is the most builder-crowded category in the Stripe Index at 34.7% micro-SaaS, more than five times the index average. AI writing tools in TrustMRR median $190 among paying startups, close to the index but far from the trades group. And the highest-gap complaints in the AI writing categories are about price, which signals a race to the bottom.
Reviewers are also increasingly specific about quality limits. One Capterra reviewer:
“Sometimes produce hallucinated or awkward output. However, tool is evolving.” – Capterra review
The YC talk offers the counterweight: a good AI-era idea sits at the edge of what models can do, and should verticalize toward owning an outcome. An AI idea can win the rubric if it pairs with a high founder-fit score or a vertical with low density; the AI opportunity index ranks where that is still true. A generic AI writing tool usually cannot. For the revenue side, read the AI SaaS revenue reality check and AI SaaS pricing models.
If one of your candidates is a marketplace, the rubric understates its risk. Marketplaces carry the highest TrustMRR median of any category ($885 among paying startups) but only 20+ paying marketplaces are in the data, and the cold-start problem does not show up in any complaint count. Founders reviewing a two-marketplace decision on r/startups were blunt:
“Marketplaces are very hard to get off the ground as you need both customers and sellers, and attracting one depends on the other.” – r/startups
Another commenter gave the most useful comparison framing in the thread:
“Both ideas have the same disease but not the same dose.” – r/startups
Add two questions to any marketplace row: can you seed one side by hand, and what stops both sides from going direct after the first match? If you cannot answer both, drop its buildability and distribution scores to 1. More in marketplace business ideas.
Repeat founders have more advantages and a few extra ways to waste them.
“Crazy to think that I’ve spent months on one thing but my natural instinct as a builder was to work on a new project rather than sell the thing that I’d been working on.” – r/SaaS
A useful sanity check for founders with a track record: is this idea an asset you could sell, or a job you are creating? Is your SaaS an asset or a job covers that test.
Some founders already run several small products and are choosing where the next hours go. The rubric still applies, but the decision is weekly, not once. One r/SaaS founder running several products solo described the failure mode:
“Whatever I touched last tends to win, not whatever actually moves the needle.” – r/SaaS
The best reply in that thread separated the two decisions:
“Equal rotation feels fair, but it can quietly starve the product that is actually closest to compounding.” – r/SaaS
Another named the emotional driver:
“Once i noticed the urge to switch was just discomfort and not a real signal it got way easier to leave the noisy one alone.” – r/SaaS
The rule that works: pick one primary product per week using revenue momentum (the TrustMRR guide shows how to benchmark each one), and define in advance what counts as an emergency for the rest. One founder put the money version simply:
“I prioritised the one growing the fastest because it funds the other projects.” – r/SaaS
A committed choice is not a permanent one. Reopen it in exactly two situations: a pre-written kill criterion fires, or going deep reveals a bigger problem underneath the one you started with. The YC talk treats the second as the normal outcome:
“Going deep isn’t primarily a process for validating the idea you started with. It’s a way to find the better idea underneath.” – Jon Xu, Y Combinator
Do not reopen it because a new idea arrived. New ideas always arrive. Capture them, park them, and rescore the list on a fixed date. One r/indiehackers founder described the system:
“The way I combat this is to totally separate idea capture and idea filtering. They’re two separate activities.” – r/indiehackers
If the evidence does point to a change, how to find product-market fit and startup pivot examples show what a disciplined move looks like.
The same themes repeat across r/startups, r/Entrepreneur, r/SaaS and r/indiehackers. The fear of picking wrong:
“I’m stuck at the point where I don’t know which idea actually deserves my time and energy. They all seem interesting in different ways, and I’m afraid of picking the ‘wrong’ one or spreading myself too thin.” – r/Entrepreneur
The honest admission that ideas can be a distraction:
“Having too many ideas is usually a distraction, not an advantage. I have been guilty of chasing novelty instead of traction.” – r/Entrepreneur
The fix most experienced commenters converge on, a time box:
“What helped me was forcing ideas to earn my time. I’d give each idea a very small, fixed test window.” – r/Entrepreneur
“Set a date and choose one on that day. It doesn’t really matter which. Go all in on making it work, give yourself a deadline say 3-6months.” – r/Entrepreneur
The boredom test:
“Which problem would I still work on even when I’m bored, tired, and not getting validation?” – r/Entrepreneur
Filters that cut a long list quickly:
“If I couldn’t describe who it’s for in one sentence, the idea wasn’t ready. killed about 80% of my list that way.” – r/indiehackers
“The real test is talking to 5 people who would actually pay for it, most of those 17 ideas would get shot down in the first conversation.” – r/indiehackers
The gut check some founders trust more than a spreadsheet:
“The question that cut through my list was not ‘which idea is best’ but ‘which one would I be embarrassed to see someone else ship next month.’” – r/indiehackers
And the warning that analysis can become its own procrastination:
“Frameworks and analysis help, but they can also become a way to feel productive without committing to anything.” – r/indiehackers
“Brainstorming ideas feels productive, but it’s really just procrastination masquerading as creativity.” – r/indiehackers
The pre-mortem question:
“I stopped asking ‘is this a good idea’ and started asking ‘what would kill this idea in month 3?’” – r/indiehackers
The defensibility question the rubric should not forget:
“How defensible is the concept from other avenues? Can it be cloned quickly and/or can future versions of AI models do it seamlessly?” – r/startups
Defensibility is not one of the nine criteria because it is hard to score before you build, but it belongs in your kill criteria. For how founders think about moats in 2026, see the SaaS moat in the AI era.
Most comparisons fail on data, not on method. A generalist assistant can build a scoring template in seconds but has no live complaint counts, saturation figures or revenue medians of its own. Ranked by how much real market evidence each can put behind a score:
| Rank | Tool | Best for |
|---|---|---|
| 1 | BigIdeasDB | All seven data criteria from one warehouse: 1M+ complaints, market-gap scores, Stripe Index density, TrustMRR revenue, funded momentum, plus the Idea Evaluator |
| 2 | ChatGPT | Turning your rubric into a template and stress-testing your assumptions |
| 3 | Claude | Writing kill criteria and interview scripts; with the BigIdeasDB MCP server it can query live data |
| 4 | Google Trends | Direction of interest in a problem term, never volume |
| 5 | Notion | Keeping the dated rubric and the parked-ideas list |
Google Trends is worth one note. Interest in “which business to start” peaks every September; its highest week in five years was in September 2025, and it sits at 53% of that peak now. “Startup ideas” is at 27% of its March 2026 peak. The indecision is seasonal. The method should not be.
All figures were computed on September 25, 2026 with read-only SQL against BigIdeasDB’s live warehouse. The three worked-example ideas were defined as groups:
Scores used the fixed thresholds in the scoring table. Founder-fit scores are for a stated hypothetical founder and are illustrative, not measured. Counts are rounded down with a plus sign; percentages, scores and medians are exact. Quotes are verbatim, attributed to subreddit or platform only, with usernames removed.
| Source | What it contributed | Limitation |
|---|---|---|
| Capterra reviews and pain points (270,000+ reviews, 39,000+ pain points) | Complaint volume, severity, market-gap scores | Coverage decays alphabetically by category; some categories have no companies. Keyword matching misses paraphrase. |
| Capterra SaaS opportunities (3,100+) | Competitive gap, feasibility | Scores are model-generated from review evidence, not market sizing. |
| G2 insights (9,000+) | Complaint volume cross-check | Smaller corpus; skews to mid-market software buyers. |
| Reddit pain points (2,000+) | Complaint volume and impact rating | Small, curated set; impact is a model rating. |
| Stripe Index (30,000+ companies) | Micro-SaaS density, buildability | Only companies on Stripe’s public directory. Classifications are AI-generated. No revenue. |
| TrustMRR (8,600+ startups) | Median MRR, $1,000 share, base rates | Skews to indie products. Keyword groups can be small; trades had under 20 paying startups. |
| Funded DB (17,000+ companies) | Momentum score | Momentum is AI-scored. Funding amounts are not used. |
| Reddit threads and the YC talks | Founder language, decision patterns | Anecdote and expert opinion, not measurement. |
| Google Trends | Seasonality of indecision searches | Relative index only, never a search volume. |
The trades revenue figure is the weakest number on this page. Fewer than 20 paying startups in TrustMRR matched the trades keywords, so its $500 median could move a lot with a few more products. That is why we capped its revenue score at 4. The direction is supported elsewhere, by low builder density and a high share of high-severity complaints, but treat the exact median as a hint.
The Stripe Index category for home services is full of operators, such as contractors and cleaners who take payments, rather than software for them. That is what makes the density figure meaningful, but it also means the category’s average buildability (3.34 across all companies) describes trades businesses, not trades software. We used the micro-SaaS subset instead.
Keyword grouping is blunt. “Invoice” appears in many products that are not bookkeeping tools, and “content” appears in many that are not AI writing tools. We checked samples by hand and the groups read correctly, but every figure should be read as a category signal, not a precise census. The rubric is a way to structure a decision, and the scores are only as good as the thresholds you set before looking.
BigIdeasDB puts all seven data criteria in one place: 1M+ complaints, Capterra market-gap scores, micro-SaaS density across 30,000+ Stripe companies, revenue medians from 8,600+ verified startups and momentum across 17,000+ funded companies. Run each candidate through the Idea Evaluator, then pull the numbers behind the score.
Open the Idea Evaluator →We built this rubric around our own data because the hardest part of comparing ideas is getting numbers you did not make up. Here is the order we would use the product in:
For the surrounding decisions: how to find problems worth solving, how to research market size, choosing an ideal customer profile, competitive landscape analysis, what each vertical can afford, and getting your first 100 users once you have chosen. Plans are on the pricing page.
Score every idea on the same nine criteria: complaint volume, complaint severity, documented market gap, saturation, revenue proof, capital flow, buildability, unfair advantage and distribution access. Weight revenue proof, unfair advantage and distribution access highest. Pick the top score, then commit to a two-week go-deep sprint with kill criteria written in advance. If two ideas finish within about 10% of each other, pick the one whose weakest score you can change fastest.
Equal-looking ideas are usually unscored ideas. Once you pull real numbers, differences appear fast: in our worked example, complaint volume differed by more than 30x between three plausible SaaS ideas, and small-software density ranged from 0.4% to 34.7% of companies in the category. If they still tie, YC's advice applies: pick one, burn the other boats, and let depth surface the better idea.
Usually no. YC General Partner Jon Xu argues that juggling ideas produces bad data, because you never go deep enough on any of them to learn whether it works. The exception is a short, time-boxed disconfirmation test of one or two days per idea before you commit. Once you commit, stop the others.
Passion is a tie-breaker, not a selector. Founders on Reddit report that almost every idea becomes boring once execution starts, so the better question is which problem you would still work on when you are tired and not getting validation. Score interest inside the unfair-advantage criterion rather than letting it override the data.
Write the deadline and the kill criteria before you start. Two weeks is enough to know whether strangers will talk to you and whether anyone reaches for a wallet. Three to six months is a common commitment window for building once that first signal exists. Switching without hitting a pre-written kill criterion is usually boredom, not evidence.
Revenue proof and distribution access. Revenue proof tells you whether anyone already pays in that category. Distribution access tells you whether you can reach buyers without paid ads. An idea that scores zero on either should be eliminated regardless of how well it scores elsewhere.
No. Competition is evidence that people pay. What matters is density: how much small software already serves the same buyer. In the Stripe Index, AI tools carry 330+ micro-SaaS products among 950+ companies (34.7%), while home services and trades carry fewer than 5 among 960+ (0.4%). Both are competitive. Only one is crowded with builders like you.
Use the same rubric but read revenue proof carefully. Consumer categories often show more startups earning something and lower medians per product. Compare the median monthly revenue of paying startups and the share that reach $1,000 MRR, not the count of startups, because a big count mostly measures how easy the category is to enter.
Check whether the low scores are about the market or about you. Market structure, such as saturation and category revenue, is hard to change. Founder fit, such as domain knowledge and distribution, can be built by going deep. If the idea loses on market structure, drop it or narrow it. If it loses on founder fit, you can choose to close that gap deliberately.
Yes, and it should carry real weight. Previous customers, an audience and knowledge of how a buyer purchases are the two founder-fit criteria in this rubric, weighted three times. But YC warns that second-time founders often weaponize founder-market fit against themselves by assuming they need a decade of domain experience. Depth can be built quickly by talking to customers.
Score the data criteria from sources you did not create, such as complaint counts, revenue medians and saturation figures, before you score founder fit. Have someone else score the founder criteria blind. Write your kill criteria first. Founders on r/startups admit that when they are excited they quietly bump pain to a 5 and reach to a 4.
Only if everything else ties. Buildability separated our three worked-example ideas by less than 0.2 points on a 10-point scale, because modern tools make most SaaS buildable. The hard part is demand and distribution, and AI makes it easier than ever to ship something nobody asked for.
Not automatically, but they start with a structural handicap on density. AI tools are the most builder-crowded category in the Stripe Index at 34.7% micro-SaaS, and AI writing tools we measured median $190 MRR among paying startups. An AI idea needs a stronger unfair advantage or distribution score to win the rubric.
It is YC's instruction to explicitly foreclose your other options once you choose: stop working on them, tell any customers you have pivoted, and focus on the chosen idea. Keep a parked list if you must, but do not touch it until a pre-written review date.
When a kill criterion you wrote in advance fires, or when going deep reveals a bigger structural problem underneath your original idea. YC's own observation is that depth often surfaces a better idea you could not see at the start. Reopen on evidence, not on a new shiny idea.
BigIdeasDB is built for this. The Idea Evaluator scores an idea's market potential, competition and feasibility in about 30 seconds, and the underlying data covers 1M+ complaints, 30,000+ Stripe directory companies, 8,600+ revenue-verified startups and 17,000+ funded companies. Generalist assistants like ChatGPT or Claude can help structure your thinking but have no live market data of their own.
From BigIdeasDB's warehouse, queried read-only on September 25, 2026: Capterra reviews and pain points, G2 insights, Reddit pain points, Upwork job pain points, the Stripe Index, TrustMRR revenue data and the funded company database, within a complaint corpus of 1M+ records. Founder quotes come from public Reddit threads, attributed to subreddit only.
BigIdeasDB Research. (2026). How to Choose Between Startup Ideas When You Only Have One Build Slot. BigIdeasDB. Retrieved from https://bigideasdb.com/how-to-choose-between-startup-ideas