Every other prompt list asks the model to imagine an answer. These ask it to reason over evidence you paste in, which is the only difference that changes the output.
BigIdeasDB is the product-research and idea-validation platform that turns documented complaints into structured demand evidence. These 22 prompts were built and tested against that data: a corpus of 1M+ complaints, reviews and discussions plus verified revenue for thousands of shipped products, so the prompts push toward what people actually pay for instead of toward what sounds good.
Every prompt list you can find asks the model to invent. Act as a skeptical investor. Generate 20 business ideas in this industry. Those prompts feel productive and produce the same output for everyone who runs them, for a reason this article can prove with revenue data rather than assert.
The best prompt for finding a business idea is one that gives the model evidence and forbids it from adding anything else. Paste in 20 to 50 real complaints from your target market, tell it to cluster and rank them using only that text, and ask it to flag every case where someone describes a workaround they built themselves. Those workarounds are people already paying in time, which is the closest thing to proof you can get before talking to anyone.
Short answer: stop asking AI to generate ideas and start asking it to process evidence. The three prompts that do most of the work are the complaint pattern extractor, the sledgehammer, and the conviction score. Add one line to every prompt you write: "if you cannot answer from the evidence provided, say so instead of guessing."
A prompt with no evidence in it is a request for the statistical centre of the model's training data. That centre is identical for every person who asks, which is why the outputs are identical. One person in r/ChatGPTPromptGenius described the experience exactly:
“every time i asked chatgpt/claude for underrated ways to make money it gave me the same tired stuff (etsy, dropshipping, faceless youtube) that's clearly already saturated to death.” - r/ChatGPTPromptGenius
The same complaint appears from people using AI for planning rather than ideation:
“Most people ask AI one giant question, like write my business plan. They get back something generic. It sounds like nobody.” - r/smallbusiness
“Small and specific beats big and vague. Every time.” - r/smallbusiness
And from founders who tested the tools specifically for ideation:
“Wouldn't entrust the idea part to AI as they only generate common ideas.” - r/Entrepreneur
The highest-voted insight in that community is not about a better prompt template. It is about changing what you ask for:
“i stopped asking Claude for answers. i started asking for frameworks. everything changed.” - r/ChatGPTPromptGenius
“don't give me an answer. give me the framework i should use to find the answer.” - r/ChatGPTPromptGenius
That is the shape of everything below. For the wider version of this argument, see our guide to finding business ideas with AI and real market problems and how to brainstorm business ideas without a model at all.
This is usually asserted. It can be measured. BigIdeasDB tracks verified revenue for thousands of shipped software products and groups them into clusters. The largest cluster is the zero-revenue cohort: 4,900+ companies with a median MRR of $0. What matters is not the size, it is the contents. The defining products of that cluster are chatbot trainers, document extraction tools, reply assistants, workflow automations, scheduling clones, Slack apps and file converters.
Read that list again next to the output of a generic idea prompt. They are the same list. Thousands of people asked a model for an idea, got the statistical average, built it, and landed in a cohort with a median revenue of zero. The technology stack is as uniform as the ideas: 29% of that cluster runs Next.js and 14% runs Supabase, which tells you how little differentiation there is anywhere in the stack.
The saturation is visible from the payments side too. The Stripe Index shows 950+ companies in the AI tools category, of which 330+ are micro-SaaS and 150+ are agentic products. More detail lives in SaaS market saturation in 2026 and the Stripe Index database.
If a generic prompt hands you an idea from the saturated centre, this is the revenue distribution you are joining. Median, not average, because the average is distorted by a handful of outliers in every cluster.
| Cluster | Companies | Median MRR | Avg growth | What it tells you |
|---|---|---|---|---|
| Zero MRR / pre-launch | 4,900+ | $0 | +34.8% | The single largest cluster we track |
| Micro-SaaS under $500 MRR | 2,500+ | $50 | +221% | Growth is real but the base is tiny |
| Flat / stable growth | 3,000+ | $0 | +0.4% | Shipped, live, and going nowhere |
| Declining | 2,200+ | $38 | -46.8% | Shrinking, and many are for sale |
| Pure SaaS / recurring | 3,300+ | $0 | +39.7% | Even the recurring cohort medians at zero |
Two rows deserve attention. The micro-SaaS cluster shows +221% average growth on a $50 median MRR, which is what growth looks like when it starts from almost nothing. And the declining cluster holds 2,200+ companies shrinking at -46.8%, many of them listed for sale. Acquisition listings tell the same story in detail: one AI resume tool reports $30,000 trailing revenue from 5 paying customers, and one consumer AI app reports $106,000 trailing revenue with a single paying customer disclosed. Those are not businesses, they are experiments with revenue attached.
The cost of skipping the evidence step is not measured in wasted prompts. It is measured in years:
“Lost runway after 4 years building... considering freelancing but risk losing momentum.” - r/EntrepreneurRideAlong
A founder in r/microsaas put the lesson in one line:
“Most founders don't fail because they can't build. They fail because they build before validating the math.” - r/microsaas
CB Insights has found the same thing at scale for years: analysis of startup post-mortems consistently puts no market need at roughly 42% of failures, ahead of running out of cash. See why startups fail in 2026 and the failure statistics for the full picture.
Every prompt in this library follows the same four-part shape. Learn the shape and you can write your own.
The constraint layer matters more than any wording trick. Both OpenAI and Anthropic give the same advice in their own prompt-engineering guidance: be explicit, use clear delimiters around supplied content, and state the output format you want.
To show why this matters, consider what changes when only the evidence changes.
Without evidence: "Give me 10 SaaS business ideas for small businesses." The output will contain an invoicing tool, a scheduling tool, a review-collection tool, a chatbot and a social media scheduler. Everyone gets that list. Cross-reference it against the zero-revenue cluster above and you will find every one of them.
With evidence: the same request, preceded by 40 real complaints from a specific software category and a constraint banning the saturated categories. Now the model is clustering text that only you have, and the output is bounded by what people actually said. It might return something narrow and unglamorous, like a tool that reconciles one export format between two systems that never talk to each other. That is what a real opportunity looks like, and it is the pattern behind most of the entries in unique business ideas backed by real complaints and business ideas that solve real problems.
Abstract advice about prompting is easy to nod along to and hard to act on, so here is the concrete version. We ran prompt 1 over a real set of complaints about business-planning software. The instruction was to cluster, count, rank by repetition and reported cost, and flag workarounds. Nothing else. The output was a table whose top rows looked like this.
| Problem cluster | Severity | Reported cost | Workaround described? |
|---|---|---|---|
| Exported plan formats inconsistently | 4.5/5 | Up to 6 hours reformatting | Yes, manual layout rebuild |
| Financials not linked to the rest of the plan | 4.5/5 | About 2 hours a week | Yes, separate spreadsheet |
| No forecasting tools at all | 4.5/5 | 4 to 5 hours a week | Yes, manual calculations |
| No real-time collaboration with a partner | 4.0/5 | Days of delay finalising | Yes, emailed versions |
| Rigid customisation of output | 3.8/5 | About 3 hours per presentation | Partially, re-export and edit |
Three things are worth noticing about that table, because they are what you are aiming for every time you run these prompts.
Every row has a cost attached. Not "users find formatting frustrating" but six hours. A problem without a number is a preference, and preferences do not become businesses. The number also gives you your pricing anchor later: if a problem costs someone four hours a week, you know roughly what solving it is worth to them.
Every row says whether a workaround exists. All five here do, which is why all five are worth a conversation. A complaint with no workaround usually means the person is annoyed but not annoyed enough to act, and that is the difference between a review and a market.
Nothing in the table was invented. Each row traces back to text someone actually wrote. Run the same prompt with no evidence block and you get a plausible list of generic software complaints instead, which reads almost identically and is worth nothing. That gap, between output that looks the same and output that is grounded, is the entire reason this library exists. The same discipline runs through finding ideas from reviews and complaints and the complaint database comparison.
The prompts are useless without input, so this is the step to get right. Real complaint text lives in software review sites, app-store reviews, support forums, community threads and your own inbox. What you want is people describing a problem in their own words, not a summary of a market.
It is also the step people most want to skip, which is why the recurring question in owner communities is whether something will just do it for you:
“I know I can prompt ChatGPT manually, but I'm wondering: are there any AI tools or platforms that already do this well out of the box?” - r/smallbusiness
BigIdeasDB exists to make that step take minutes instead of a week. It aggregates and severity-scores complaints across 11-plus sources into a searchable corpus, so you can pull a category's recurring problems in one query. Start at Discover, or go straight to pain points and complaints. The business idea generator and idea evaluator are free entry points, and tools to find customer pain points covers the alternatives honestly.
For method rather than tooling, see finding ideas from real user pain points, mining Capterra reviews, turning G2 reviews into ideas, and negative reviews as a source.
Work through them in order the first time. After that, most people use four or five repeatedly. The right-hand column is the part that decides whether the prompt works.
| # | Stage | Prompt | What it does | What you paste in |
|---|---|---|---|---|
| 1 | Find the problem | Complaint pattern extractor | Turns raw complaint text into ranked, repeating problems | 20 to 50 real complaints |
| 2 | Find the problem | Workaround detector | Finds problems people already pay or work to solve | The same complaint set |
| 3 | Find the problem | Budget-holder locator | Separates the person in pain from the person who pays | One problem cluster |
| 4 | Find the problem | Trivial-versus-purchasable filter | Kills complaints nobody would pay to fix | Your ranked problem list |
| 5 | Shape the idea | Constrained idea generator | Generates ideas bounded by your evidence, not by its training data | Your purchasable problems |
| 6 | Shape the idea | Narrowing prompt | Turns a broad idea into a wedge | One idea |
| 7 | Shape the idea | Unfair advantage mapper | Finds the version only you can build | Your own background |
| 8 | Shape the idea | Existing-solution check | Stops you reinventing something that already ships | The idea plus what you found searching |
| 9 | Pressure-test | The sledgehammer | Attacks the idea instead of encouraging it | The idea and your evidence |
| 10 | Pressure-test | Assumption ranker | Orders your assumptions by what would kill you fastest | The idea |
| 11 | Pressure-test | Pre-mortem | Writes the failure story before it happens | The idea |
| 12 | Pressure-test | Steelman the do-nothing option | Tests whether the customer will actually switch | The idea plus the current workaround |
| 13 | Size it | Bottom-up sizing | Builds a market number you can defend | Your buyer definition |
| 14 | Size it | Sanity-check against reality | Compares your forecast with what similar businesses actually earn | Real benchmark data |
| 15 | Size it | Saturation reader | Tells you whether crowded means dead or means proven | Competitor and market counts |
| 16 | Buyer and price | Discovery script builder | Turns evidence into questions that get honest answers | Your problem cluster |
| 17 | Buyer and price | Willingness-to-pay probe | Finds the real price, not the polite one | Interview notes |
| 18 | Buyer and price | Objection generator | Surfaces the sales objections before your first call | The offer |
| 19 | Buyer and price | Channel reality check | Kills channels you cannot actually work | Buyer plus your honest constraints |
| 20 | Decide | Conviction score | Forces a number and a next action | Everything above |
| 21 | Decide | Kill-or-continue gate | Makes the decision explicit | The score and your tests |
| 22 | Decide | First-week plan | Converts the decision into work | The surviving idea |
Nothing downstream works if this stage is skipped. You are looking for problems that repeat, cost something measurable, and already have someone working around them by hand.
Below are real complaints from users of [CATEGORY] software, copied verbatim. <complaints> [PASTE 20-50 REAL COMPLAINTS HERE] </complaints> Do not add problems that are not in this text. Working only from what is above: 1. Cluster these into distinct problems. Name each one in under 8 words. 2. For each cluster, count how many complaints it covers and quote the two most specific lines. 3. Rank clusters by how often they repeat AND how much time or money the complainant says it costs. 4. Flag any cluster where people describe a workaround they built themselves. Those are the strongest. 5. List which clusters are about the product versus about pricing, support, or onboarding. Output a table. No recommendations yet.
Using only the complaints below, find every instance where someone describes a manual workaround: a spreadsheet, a second tool, a virtual assistant, a script, or a process they invented. <complaints> [PASTE COMPLAINTS HERE] </complaints> For each workaround: quote it, state what it replaces, and estimate the effort it costs based only on what the person said. Rank by effort. A documented workaround is proof someone is already paying in time. Ignore complaints with no workaround.
Problem: [PASTE ONE PROBLEM CLUSTER AND ITS SUPPORTING QUOTES] Answer these separately and do not merge them: 1. Who experiences this problem day to day? Give the job title. 2. Who controls the budget that would pay to fix it? Give the job title. 3. If those are different people, what does the budget holder care about that the user does not? 4. What is the internal justification the budget holder would have to make to approve this? 5. Where do both of these people already spend money in this category? If you cannot answer from the evidence provided, say "not determinable from this evidence" instead of guessing.
Here is a ranked list of problems: [PASTE LIST] For each one, classify it as PURCHASABLE or TRIVIAL using these tests only: - Does someone lose measurable time or money? (quote the evidence) - Has anyone described building a workaround? - Does it recur, or is it a one-off annoyance? - Would it still matter if the person had 10x more of it to do? A problem is TRIVIAL unless it passes at least three. Be strict. Most complaints are trivial. Explain each verdict in one sentence.
Prompt 2 is the one to run if you only run one. A documented workaround is the strongest pre-conversation signal there is, because someone has already decided the problem is worth their labour. The logic is the same as validating demand with Upwork jobs: people paying to solve something is better evidence than people saying they would.
Now you convert problems into candidate products. Prompt 5 carries the explicit ban on saturated categories, which is the single most important line in this whole library.
You will propose product ideas, with hard constraints. Evidence: [PASTE YOUR PURCHASABLE PROBLEM LIST WITH QUOTES] Constraints: - Every idea must trace to a specific quoted complaint above. Cite the quote. - Do not propose anything that requires a marketplace, a two-sided network, or venture funding. - Do not propose an AI chatbot, an AI writing tool, a resume tool, a Calendly clone, or a document converter. These are saturated. - Each idea must be shippable by one person in under 8 weeks. Give 6 ideas. For each: the quote it comes from, the single job it does, and who pays. Reject your own first three ideas as too obvious and replace them before answering.
Idea: [PASTE IDEA] Narrow it five times. Each time, cut the scope in half by picking a smaller audience or a smaller job. Show all five versions. Then tell me which version could get its first paying customer fastest, and why. Do not tell me the broadest version is best.
My background: [PASTE YOUR REAL EXPERIENCE, SKILLS, NETWORK, AND ACCESS. BE SPECIFIC AND BORING.] Candidate idea: [PASTE IDEA] 1. Which parts of this idea does my background make easier than it would be for a stranger? 2. Which parts does it not help with at all? 3. Is there a nearby version of this idea where my background matters more? 4. What would I have to learn or buy to close the gaps? Do not flatter my background. If it is irrelevant to this idea, say so.
Idea: [PASTE IDEA] What I found when I searched for existing solutions: [PASTE WHAT YOU ACTUALLY FOUND, INCLUDING NAMES AND PRICES] Based only on what I pasted: 1. Which of these already do the core job? 2. What specifically do their users complain about? Quote if I gave you quotes. 3. Is there a real gap, or am I proposing a slightly different version of something that exists? 4. If there is a gap, state it in one sentence a buyer would recognise. An existing competitor proves the market. Tell me whether the gap is real, not whether the space is occupied.
On prompt 8: an existing competitor is validation, not a kill. The question is whether the gap is real and whether you can reach the segment they ignore. That reasoning is worked through in niche viability validation and the underserved software markets study.
Models default to encouragement, which makes them dangerous at exactly this step. These four prompts are written to fight that default. One prompt author in r/ChatGPTPromptGenius framed the need well:
“Most prompts out there are just cheerleaders. This one is a sledgehammer.” - r/ChatGPTPromptGenius
“If your idea survives this, you're actually onto something. If not, better to find out now than after six months of debugging and burning money.” - r/ChatGPTPromptGenius
You are reviewing this idea to find reasons it will fail. You are not here to be encouraging, and you will not end with a positive summary. Idea: [PASTE IDEA] Evidence I have: [PASTE EVIDENCE] 1. What is the single assumption this entire idea rests on? 2. What is the most likely reason this fails in the first 6 months? 3. What am I assuming about the customer that I have not verified? 4. Which of my evidence points are weak or circular? 5. What would a person who has already tried this tell me? End with the one thing I should go find out this week. No summary, no encouragement.
Idea: [PASTE IDEA] List every assumption I am making about the customer, the problem, the market, distribution, pricing, and my own ability to execute. Then rank them by: how likely each is to be wrong, multiplied by how badly the business breaks if it is wrong. For the top three, give me the cheapest test that would tell me the truth in under a week, with a specific pass or fail threshold. A test with no threshold is not a test.
It is 18 months from now and this business has failed: [PASTE IDEA] Write the post-mortem as the founder. Be specific about what happened month by month. Name the moment it became unrecoverable and the earlier decision that caused it. Then list the three signals that were visible in month two that would have predicted this.
Idea: [PASTE IDEA] What the customer does today instead: [PASTE THE CURRENT WORKAROUND] Argue as convincingly as you can that the customer should keep doing what they do today and not buy this. Cover switching cost, the risk of change, who gets blamed if it fails, and whether the problem is painful enough to bother. Then tell me what would have to be true for the switch to happen anyway.
Prompt 12 is underrated. Most ideas do not die because a competitor is better, they die because the customer keeps doing what they already do. The same failure mode drives most of SaaS churn and shows up throughout failed business idea lessons and common validation pitfalls.
This is where AI output is least trustworthy, so the prompts are built to force labelling and to refuse when inputs are missing.
Buyer: [PASTE THE SPECIFIC BUYER DEFINITION] What I know about how many exist: [PASTE ANY REAL COUNTS YOU HAVE] Build a bottom-up estimate: number of buyers, times realistic annual price, times a conservative reachable share. Show every step. Label each input as MEASURED (I gave you a real number), ESTIMATED (you derived it, say how), or UNKNOWN. Do not produce a top-down market size. Do not cite an industry report. If the UNKNOWN inputs dominate, say the estimate is not usable yet and tell me which single number to go and find.
My revenue projection: [PASTE PROJECTION] Real benchmark data for comparable businesses: [PASTE ACTUAL MRR, MARGIN AND GROWTH FIGURES YOU LOOKED UP] Compare my projection to the benchmarks. Where am I assuming I will be an outlier? Quantify how much of an outlier. If my projection sits above the top quartile of the benchmarks, say so plainly and tell me what would have to be exceptional about my execution to justify it.
Category: [PASTE CATEGORY] What I found: [PASTE NUMBER OF EXISTING PLAYERS, THEIR PRICES, AND ANY REVENUE DATA] Answer: 1. Is this category crowded with operators, or crowded with software? Those are different. 2. Where the incumbents are old or badly reviewed, is that an opening or a warning? 3. What would a new entrant have to be 10x better at to matter here? 4. Is there a segment inside this category that the existing players ignore? Treat existing competition as validation of demand. Only call it a kill if there is a dominant, well-liked incumbent serving my exact segment.
For prompt 14 you need real benchmarks to paste in. Useful anchors: across the AI category we track, average MRR is about $1,410 with a median of $0 and margins near 65%; in marketing tools, 483 tracked startups average $1,785 with a median of $0 and 99 currently for sale. Full method in how to calculate market size, revenue benchmarks by category, SaaS metrics benchmarks and how fast startups actually grow.
Four prompts that convert an idea into a sales hypothesis you can test on a call this week.
Problem: [PASTE PROBLEM AND QUOTES] Write 10 customer discovery questions I can ask on a 20-minute call. Rules: - Ask only about what they did in the past, never about what they would do in future. - Never mention my idea or any solution. - No yes-or-no questions. - At least two questions must uncover what they currently spend on this, in money or hours. For each question, tell me what a bad answer looks like so I know when I am being told what I want to hear.
Here is what buyers told me: [PASTE YOUR REAL INTERVIEW NOTES] Working only from these notes: 1. What did they say the current problem costs them? Quote it. 2. What are they already paying for adjacent tools? 3. What price would be an obvious yes, and what price would need approval? 4. Where in these notes did someone signal real intent versus being polite? Quote both. Then propose three price points and tell me which one I can defend from the evidence above. If the evidence does not support a price, say so.
My offer: [PASTE THE ONE-SENTENCE OFFER AND PRICE] Buyer: [PASTE BUYER] List the 8 objections this buyer will raise, ordered by how often they will come up. For each, give the real underlying concern, not the stated one. Then give me one honest response to each. Do not give me a way to dodge the objection. If an objection is fair and I cannot answer it, say that.
Buyer: [PASTE BUYER] My real constraints: [PASTE YOUR BUDGET, YOUR WEEKLY HOURS, YOUR EXISTING AUDIENCE OR LACK OF ONE] List the channels where this buyer can be reached. For each: the realistic cost to get one customer, the time before the first result, and whether it works at my constraints. Cross out every channel that requires money or an audience I do not have. Of what remains, tell me the one to start with and what week one looks like.
Prompt 16 exists because badly worded discovery questions produce agreeable, useless answers. Pair it with the full customer discovery questions list. Prompt 19 exists because founders repeatedly pick channels they cannot work, a failure described bluntly in a community thread:
“I posted on x, linkedin outbound, promos on reddit, forced users through a 10 step onboarding without knowing retention... stuck at 0.” - r/EntrepreneurRideAlong
“We wasted weeks obsessing over growth hacks. What actually worked was boring, manual work: commenting on reddit, directories, founder stories on linkedin.” - r/EntrepreneurRideAlong
“Stop selling generic software. Niche down and sell solutions for a specific business type, get testimonials and case studies before scaling.” - r/EntrepreneurRideAlong
More on that in getting your first customer, getting customers for a startup, and pricing strategies.
The last three prompts force a number, a kill condition and a week of work. Without them the whole exercise stays a nice conversation.
Score this idea from 0 to 10 on how convinced I should be, using only the evidence below. Evidence: [PASTE EVERYTHING: COMPLAINTS, INTERVIEWS, BENCHMARKS, PRICING NOTES] Give: the score, the single assumption holding the score up, the two pieces of evidence that most support it, the two that most undermine it, and the cheapest test that would move the score by two points in either direction. If the evidence is thin, score it low. A low score with a clear next test is more useful to me than a high score.
Idea: [PASTE IDEA]. Current conviction score: [PASTE SCORE]. Tests I have run: [PASTE RESULTS] Write the specific, measurable conditions under which I should kill this idea, and the conditions under which I should continue. Use numbers and dates. Then tell me which condition I am closest to hitting right now.
Idea: [PASTE SURVIVING IDEA]. Riskiest assumption: [PASTE IT] Plan the next 7 days to test that one assumption and nothing else. Give me a day-by-day list with a specific output for each day. No building. No design. No brand work. If the plan includes writing code, replace it with something cheaper that produces the same information.
The scoring idea is not ours. A prompt author described the goal precisely:
“No vibes. A number, a reason, and the cheapest test that moves the score.” - r/ChatGPTPromptGenius
Then go and run the test. Our 8-stage validation framework, multi-signal validation and validating before you write code cover what the tests should be.
Different steps reward different things. This is not brand loyalty, it is matching a capability to a job.
| Step | What to reach for | Why |
|---|---|---|
| Extracting patterns from pasted evidence | Long-context assistant | Holds the whole complaint set without dropping rows |
| Adversarial review of your own idea | A different model from the one that drafted | It does not share the drafting model's assumptions |
| Anything needing a public source | A search-grounded assistant | You need a link you can check, not a confident sentence |
| Numbers and sizing | A spreadsheet, then AI to attack it | Generated projections drift toward the wrong business model |
| Turning notes into readable prose | Any general assistant | This is the one job all of them genuinely do well |
Founders who use these tools heavily describe a mixed stack rather than one favourite:
“I use ChatGPT for general research, Claude for the creative aspect, and Perplexity for deep research. Most of the work is still by me and most of the idea is still by me.” - r/Entrepreneur
Full scoring of the assistants for founder work is in best AI for business planning and best AI tools for entrepreneurs.
A model is poor at finding faults in text it just produced. The fix is mechanical: move the output to a different assistant before running the pressure-test prompts.
“Copy and paste the convo from ChatGPT into Claude or Gemini... to challenge the idea and identify weak points and gaps.” - r/Entrepreneur
A workable chain: extract patterns in model A, shape ideas in model A, paste the result into model B for prompts 9 through 12, then bring both outputs back and decide yourself. The deciding step is not delegable, which is the whole point of keeping delusional thinking out of validation.
Put this in your custom instructions or project settings and every prompt above gets better without editing it:
When I give you evidence, reason only from that evidence. Never invent statistics, market sizes, company names, or quotes. If a number is not in what I gave you, either omit it or label it clearly as your estimate and show how you derived it. If you cannot answer from the evidence provided, say "not determinable from this evidence" rather than producing a plausible answer. Do not end responses with encouragement or a positive summary. Do not tell me an idea is promising unless you cite the specific evidence that makes it promising. Prefer specific and narrow over broad and impressive.
The last two lines do the heavy lifting. Everything else in this article is an attempt to work around a model's default toward agreeableness, and this states it directly.
Confident, wrong output is the documented failure mode of these tools, not an edge case. In our analysis of AI assistant apps, the most severe recurring user complaint across products rated between 4.1 and 4.3 stars was inaccurate or unhelpful information, classified critical or high in every case. Popular does not mean reliable. Watch for these tells:
Three edits cover most cases. First, change the banned-category list in prompt 5 to whatever is saturated in your specific space, which you find by looking at what already exists rather than guessing. Second, change the output format demand to match how you actually work, since a table you will not read is worse than a ranked list you will. Third, tighten the constraint block every time the model wriggles out of one, because the constraint layer is where the quality lives.
If you do not yet have a market to point these at, work backwards from demand instead of from interest: how to find a profitable niche and how to find startup ideas both start from evidence rather than brainstorming, and the opportunity score explains how we rank one lead against another.
If you are working in a specific vertical, the industry-level gaps are mapped in niche SaaS opportunities by industry, most profitable SaaS niches and boring industries begging for micro-SaaS.
The prompts you use repeatedly should not live in your chat history. This is a real and documented gap: in our app-store analysis, the ability to save and reuse custom prompts appears as an explicit missing feature request on AI assistant products, with medium demand intensity. Keep them in a notes file or a project workspace, versioned, with the evidence block separated from the instruction so you can swap markets without rewriting the prompt.
Keep the evidence versioned too. The complaint set you pull today will look different in six months, and the difference is itself a signal. That is the same reasoning behind refreshing category research in the most-hated software study and the most-requested features report rather than treating either as settled.
The second mistake is the expensive one. Run the pressure-test prompts on a candidate before you have told anyone about it, and use a validation tool or the documented pain-point set to check the answer rather than trusting the model twice.
Three things stay yours. Deciding whether the problem is real, because only buyers can settle that. Producing the market size, because a generated number is the first thing anyone will test. And the final judgment on whether to continue, because encouragement is the model's default setting. A founder summarising what stayed overhyped after a year of testing put it plainly:
“What's still overhyped: fully autonomous AI agents, set it and forget it automation, AI replacing your whole team.” - r/smallbusiness
“It's useful for overcoming writer's block, structuring sections, cleaning up your wording. The downside is that it stays pretty generic.” - r/smallbusiness
Users of AI planning products describe hitting the same ceiling from the other direction:
“If you're looking for deep market analysis or a super flexible editor, you'll still want to polish things up after exporting.” - Capterra review, AI business plan software
That is an accurate description of the tool. Structure and speed, not judgment. See what to do when you have no ideas and how to decide what business to start for the human half.
If the idea survives that week, the next steps are turning it into a startup and refining the idea itself. If it does not survive, that is the prompts working.
One founder's reason for sequencing it this way is worth keeping in mind:
“I'd rather not deal with spending a bunch of time trying to find investors until I've verified that the business is a good idea.” - r/Entrepreneur
Evidence is separated into layers so no figure is overstated. The 1M+ corpus figure is historical and cumulative and is never summed with the snapshots below it.
| Source | Records | Evidence | Limitation |
|---|---|---|---|
| Complaint corpus | 1M+ | Cross-source historical record | Historical and cumulative, never a live count |
| Capterra pain points | 39,000+ | Severity-scored software problems | AI-extracted subset, not every review |
| TrustMRR revenue clusters | 5 clusters | What shipped products actually earn | Self-reported, heavily pre-revenue |
| Zero-MRR cluster | 4,900+ companies | The saturation evidence | Membership is a snapshot, not a fate |
| Acquisition listings | 8 AI businesses | Real prices and customer counts | Seller-reported, unaudited |
| Stripe Index AI category | 950+ companies | Operator density | Counts operators, not revenue |
| App-store AI assistants | 6 analysed | How the models fail in production | Small sample of one app genre |
| Upwork demand signals | 7 patterns | What people pay to have researched | Frequency only, budget fields are empty |
Source documentation lives under data sources overview, pain point analysis, complaint search and SaaS opportunities.
The 22 prompts were written against the failure patterns documented in this article rather than collected from other prompt lists. Each one was designed around a single rule: it should produce a visibly worse answer when you give it no evidence, so the prompt fails loudly instead of inventing plausible output. Where a prompt could either accept an assertion or demand a source, it demands the source.
The revenue clusters come from BigIdeasDB's verified-revenue dataset, snapshot September 2026, grouped by revenue tier, business model and growth pattern. Acquisition figures come from live marketplace listings. Saturation context comes from the Stripe Index. Community material is used for voice and pattern identification only, is quoted anonymously by subreddit with no usernames or post identifiers, and is never treated as a statistic.
The limits matter. Revenue figures are self-reported by founders, unaudited, and skew heavily pre-revenue, which is why every table here reports medians rather than leaning on averages. Acquisition listings are seller-reported and several show internally inconsistent metrics, so they illustrate rather than prove. Cluster membership is a snapshot, not a verdict on any individual company, and products move between clusters. The app-store analysis covers a small sample of one genre. And no prompt, however well constructed, substitutes for five conversations with real buyers. Read the complaint analysis platform and the idea validation walkthrough for how the underlying data is produced, or the market research guide for the non-AI version of this whole process. If you would rather start from a curated list than a blank prompt, try SaaS ideas backed by pain points, low-competition SaaS ideas or one-person business ideas.
The best prompt is not a prompt for generating ideas at all. It is a prompt for extracting patterns from evidence you paste in. Asking a model to invent business ideas returns the same handful of saturated concepts to everyone who asks, because it is drawing on the same training data every time. Asking it to cluster and rank 30 real complaints you collected returns something specific to your market. The complaint pattern extractor at the top of this list is the one to start with.
Because a generic prompt gives it nothing to work with except the statistical centre of its training data, and that centre is the same for every user. The evidence shows up in the revenue data: the largest cluster of startups BigIdeasDB tracks is the zero-revenue cohort at 4,900+ companies, and its defining products are exactly the things models suggest, including chatbot builders, document converters, reply assistants and scheduling clones. Same prompt, same idea, same crowded market. The fix is to constrain the prompt with evidence and with an explicit ban on the saturated categories.
Give the model your actual evidence, tell it not to add anything that is not in the evidence, and ask for a verdict with a threshold rather than an opinion. The three prompts that do most of the work here are the sledgehammer (find reasons this fails), the assumption ranker (order assumptions by what kills you fastest), and the conviction score (a number, the assumption holding it up, and the cheapest test that moves it). Validation still requires talking to real buyers. A prompt can only tell you what to go and ask.
Role prompts help less than people think and can hurt. Asking a model to act as a skeptical investor produces the tone of skepticism without its substance, because the model has no information about your market that it did not have before you assigned the role. What actually changes the output is constraint and evidence: what it may not say, what it must cite, and what real text it is reasoning over. Use roles for register, use constraints for rigor.
Match the model to the step rather than picking one for everything. Use a long-context assistant for extracting patterns from a large pasted complaint set, a search-grounded assistant for anything that needs a checkable public source, and a different model from your drafting one for the adversarial review, since a model is poor at finding faults in its own output. For numbers, use a spreadsheet and let AI attack it. All the major assistants are capable enough for the writing step.
Anywhere people describe a problem in their own words: software review sites, support forums, app-store reviews, community threads, and your own customer emails. BigIdeasDB aggregates these into a searchable, severity-scored corpus of 1M+ complaints so you can pull a category's recurring problems in one query instead of reading threads for a week. The important part is not the source, it is that the text is real and that you paste it in rather than asking the model to recall it.
Yes. They are written as plain instructions with explicit constraints and no model-specific syntax, so they work in any current assistant. The XML-style tags around pasted evidence in a few of them are a formatting convention that helps most models separate your data from your instructions, and they do no harm where they are not needed. Both OpenAI and Anthropic recommend clear delimiters and explicit constraints in their own prompt-engineering guidance.
Long enough to include your evidence and your constraints, which usually means the pasted material dwarfs the instruction. A one-line prompt produces a one-size-fits-everyone answer. The pattern that works is short instruction, hard constraints, large evidence block, and a demand for a specific output format. Several of the prompts here explicitly tell the model to say it cannot determine something rather than guess, which is the single most useful line to add to any prompt of your own.
You can, and for a first shortlist it is faster. The limit is the same one these prompts are designed around: a generator with no evidence layer is still producing the statistical average of its training data. Look for tools that show you the underlying complaints, reviews or revenue figures behind a suggestion, and treat anything that cannot show its evidence as a brainstorming aid rather than research.
Three things. Never let it decide whether the problem is real, because only buyers can tell you that. Never let it produce your market size, because it will generate a confident number with no source. And never let it be the last word on whether to continue, because models are trained toward encouragement and will find reasons your idea works. Use AI to structure, extract, challenge and summarise. Reserve judgment for evidence you gathered.
BigIdeasDB Research. (2026). 22 AI Prompts for Finding and Pressure-Testing a Business Idea (2026). BigIdeasDB. Retrieved from https://bigideasdb.com/ai-prompts-for-business-ideas