AI personas answer the question you ask before launch. Real customers leave over what happens after it. We measured the gap across 400,000+ reviews and the 2026 research.
Synthetic users can help you prepare for customer research. They cannot replace it. An AI persona answers the question you ask before launch: would you use this? Real customers leave over what happens after they pay. Across 9,400+ one and two star Capterra reviews, 62.1% are about support, billing, cancellation, bugs or updates. Only 9.9% name a missing feature.
That gap is the whole argument. A synthetic user has never waited three days for a support reply, never been charged after cancelling, never had an update break a workflow. So it cannot warn you about the things that actually make customers churn, and it cannot tell you whether anyone will pay. We checked the pattern against 91,000+ one and two star app reviews, the churn flags on 39,000+ documented pain points, and the 2026 academic research on synthetic participants.
This page is for founders and product people deciding whether AI personas can stand in for talking to customers. If you want the broader method of using AI for research, read how to use AI for market research. If you are about to get on calls, start with our customer discovery questions.
The rest of this page shows the numbers behind that answer, where synthetic users genuinely earn their place, and a workflow that gets you real customer voice without weeks of recruiting.
A synthetic user is an AI-generated persona you can interview. You describe a segment, such as “operations managers at 20-person logistics firms”, and a large language model answers as that person. Dedicated tools wrap this in panels, dashboards and survey exports. A general chat model like ChatGPT or Claude does the same with a role prompt.
Three things get confused under the same label, and the difference matters:
The Nielsen Norman Group put the core problem bluntly in its evaluation of synthetic users: they provide “artificial research findings produced without studying real users.” The output looks like research. It is a prediction of what research might say. For a definition of who you should be studying in the first place, see our guide to the ideal customer profile for startups.
Because the pitch got louder in 2026. On Google Trends, worldwide interest in “synthetic research” was flat for four years, stepped up in August 2025 and peaked the week of May 31, 2026. It now sits at 24% of that peak. “Synthetic users” peaked on May 17, 2026 and is at 30%. “AI personas” peaked the same week as synthetic research and is at 26%.
The most-watched YouTube talk on the subject, a July 2026 conference session on persona engineering, has 25,000+ views. NN/g’s explainer video has 12,000+. Four of the nine pages ranking for “synthetic users vs real users” are written by companies selling synthetic research, and most of the rest are UX-only. So most of what a founder reads on this topic comes from the people with the most to gain from a yes.
The founder version of the question is older than the tools. People have always wanted a shortcut past the awkward part:
“I want to test out my idea before putting money into it, but traditional market research takes too long and costs too much.” – r/solopreneurs
“How do we distinguish must-haves from nice-to-haves early?” – r/leanstartup
Synthetic users promise to answer both in an afternoon. The data below shows which half of that promise holds. If you are still choosing what to build, our guides to finding startup ideas and choosing between startup ideas come first.
Most comparisons of synthetic and real users test the simulation against a study. We took the other side: what do real customers actually talk about when a product fails them, and could a simulated customer have said it?
We ran read-only SQL over BigIdeasDB’s live tables on September 25, 2026. We split 270,000+ Capterra reviews and 136,000+ App Store and Google Play reviews by star rating, then tagged the text for themes with keyword patterns: support, billing and cancellation, bugs and outages, updates, price, missing features, integrations and ease of use. A review can match several themes. We also pulled churn flags from 39,000+ documented pain points, revenue base rates from 8,600+ revenue-verified startups on TrustMRR, and counted synthetic-research tools across the Stripe Index, funded companies, and the AI connector census.
Keyword tagging undercounts and overcounts at the edges. The full method and every limitation are in the methodology and data sources sections.
Real customers mostly complain about the relationship after purchase, not the product idea before it. In one and two star Capterra reviews, 62.1% mention support, billing, cancellation, bugs or updates. 36.6% mention only those themes and nothing about price, features, integrations or ease of use. Missing features show up in 9.9%.
A synthetic user interviewed about a concept will happily discuss features, workflows and pricing tiers, because that is what a concept interview is about. It will not bring up the support queue, the auto-renewal or the update that moved a button, because it has never used your product. The complaints that end customer relationships are the ones a simulation is structurally unable to produce.
“Everything. I have gotten zero help. Cannot login and cannot get anyone to answer my emails or call me back. I have been trying to cancel for months because no one will help me and I cannot get help to do that either! Very dissatisfied!” – Capterra review, 1 star
Nobody predicts that review in a concept test. It is still the review that decides whether the next buyer signs up, and it is the kind of evidence our business pain points research is built from.
Support is the single largest theme in negative B2B software reviews, at 42.7%. Price is 18.6%. Missing features are 9.9%. Here is the full split.
| Theme | Share of 1-2 star reviews | Can a synthetic user report it? |
|---|---|---|
| Support or customer service | 42.7% | No. Requires having needed help |
| Price or cost | 18.6% | Partly. Can guess, never pays |
| Billing, contracts, cancellation, refunds | 18.4% | No. Requires a billing history |
| Bugs, outages, broken features | 17.8% | No. Requires using the product |
| Missing features | 9.9% | Yes, and it over-produces them |
| Integrations and sync | 9.9% | Partly. Generic, not your stack |
| Updates or version changes | 8.9% | No. Requires a before and after |
| Ease of use | 8.4% | Partly. Opinion, not observed struggle |
| Any post-purchase theme | 62.1% | No |
| Only post-purchase themes | 36.6% | No |
The feature row is the one synthetic research is good at, and it is the smallest. NN/g saw the same thing from the other direction: asked what makes a course engaging, its synthetic user listed seven factors with no sense of which mattered. “Real people care about some things more than others. Synthetic users seem to care about everything,” the team wrote. To see how that priority shows up in software categories, read our state of SaaS pain points report.
“Requests for help. I submitted several tickets for help and got no response. There is no way to cancel your account by yourself. You have to submit a help ticket for this, and of course, I got no response.” – Capterra review, 1 star
“Stability issues, Support issues, Consistency issues, Accuracy issues. We spent more time working on issues than using the product.” – Capterra review, 2 stars
Consumer apps show the same shape at a different scale. In one and two star App Store and Google Play reviews, 37.5% mention billing, trials, bugs, updates, support or lost data. Missing features appear in 4.1% and ease-of-use complaints in 1.4%. That is 9 post-purchase complaints for every feature request.
| Theme | 1-2 star reviews | 3-5 star reviews |
|---|---|---|
| Bugs, crashes, not loading | 14.4% | n/a |
| Billing, trials, subscriptions, refunds | 14.2% | n/a |
| Updates or new versions | 9.5% | 8.8% |
| Support | 8.0% | n/a |
| Price, paywall, premium | 7.0% | n/a |
| Ads | 6.7% | 6.1% |
| Missing features | 4.1% | n/a |
| Leaving: cancel, uninstall, switch | 11.8% | 3.3% |
| Names a dollar amount | 5.0% | 2.0% |
Low-rated app reviews are 3.6x as likely to describe leaving as the rest. They are 2.5x as likely to name a dollar figure. That is behavior, recorded after the fact, with a number attached. For the full mobile picture, see our state of mobile app pain points.
“I downloaded the app to test out the free version and was immediately charged for a subscription upon opening the app for the first time. Now trying to figure out how to get a refund” – App Store review, 1 star
“I had a subscription for radarbox, after the update my app stopped working, had to reinstall. it does not recognize my subscription. asking to pay double for a new subscription. Basically scam.” – App Store review, 1 star
Even on a narrow match (just “support” or “customer service”), 39.1% of 1-2 star Capterra reviews raise it, against 21.7% of higher ratings. The theme almost doubles as ratings fall. It is also the theme a concept interview never touches, because nobody needs support for a product they have not bought.
“Their support is dreadful. A week or two can go by with several emails to them and all we hear are crickets. There is no response sometimes until the third or fourth email. This is terribly unprofessional and is horrible for our agency and clients.” – Capterra review, 1 star
“No Support Company seems unstable Never get to speak to the same person from support twice Seems to be forgotten after purchase” – Capterra review, 2 stars
“Poor data. Influencers often miscategorized. No response from customer support.” – G2 review
“Forgotten after purchase” is a sentence no simulated buyer writes. For a founder, it is also a wedge: in crowded categories, a better support experience is often a stronger differentiator than another feature. Our customer support software limitations report shows how often that gap recurs, and what a SaaS moat looks like in the AI era explains why service is one of the few that survive.
18.4% of 1-2 star Capterra reviews and 14.2% of 1-2 star app reviews are about how money moved: auto-renewals, charges after cancelling, trials that billed early, refunds refused. A synthetic persona can opine on a price. It cannot be surprised by a charge.
“Information was not complete and they auto renewed me without any warning and we I canceled the same day they told me to pound sand.” – Capterra review, 1 star
“They said nothing about needing a plan, then charged me for a full year membership after I selected just the monthly plan. They would not provide a refund or change the plan length.” – Capterra review, 1 star
“I haven’t used or updated the app in nearly a year, and today, without any notification or heads up of any sort, they have charged me $119.99. They are refusing to refund. This is definitely fraudulent and I will be reporting this app.” – App Store review, 1 star
“It's hard to cancel and they charge weekly. After you send the cancelation request they still take money from your acct weekly. Scam!!!” – App Store review, 1 star
If you are pricing a product, these are the design decisions that decide your reviews. Our guide to what micro SaaS actually charges covers the price points. This data covers the part pricing pages leave out. Our app store database guide shows how to pull these billing complaints for any app category.
8.9% of 1-2 star Capterra reviews and 9.5% of 1-2 star app reviews blame an update. This is the purest example of an experience that needs a before and an after. A synthetic user is always meeting your product for the first time.
“Since the recent update, everything feels messy and unnecessarily complicated. The interface is slower, harder to navigate, and many simple features are now buried or confusing to access. It’s made routine tasks more time-consuming and frustrating.” – Capterra review, 1 star
“The updates. Their timing is terrible. It is almost if they expect that you should have IT person on staff. They take no real responsibility for the actions that constantly cost their customers time and money.” – Capterra review, 2 stars
“Used to absolutely love this app. So much so that I paid for the premium version when it was a fixed rate. Now they’ve converted to a subscription and have gradually whittled down the features that I used to have.” – App Store review, 1 star
Not all pain points are equal, and the ones that predict churn are post-purchase. Across 39,000+ documented Capterra pain points, those about customer service carry a churn-risk flag 98.2% of the time and customer support 96.7%. Feature limitations carry it 57.4% of the time.
| Pain point category | Churn-risk flag | Avg severity |
|---|---|---|
| Customer service | 98.2% | 4.11 |
| Customer support | 96.7% | 4.06 |
| Pricing | 93.4% | 3.86 |
| Integration issues | 87.2% | 4.01 |
| Performance | 80.2% | 3.78 |
| Reporting | 74.8% | 3.89 |
| Functionality | 68.6% | 3.71 |
| User experience | 67.5% | 3.75 |
| Feature limitations | 57.4% | 3.65 |
This is the priority signal synthetic panels flatten. A persona lists everything as important. Real evidence ranks it, and the ranking puts the relationship above the roadmap. If you are sorting your own backlog, our guide to prioritizing feature requests from customer reviews applies the same logic, and the Capterra analysis guide shows where these churn flags live.
Angry reviews are specific. In 1-2 star Capterra reviews, 4.4% name a dollar amount, against 0.3% of other reviews, a 15x difference. 8.8% describe switching, migrating or cancelling, against 1.5%. Across all 270,000+ reviews, 14,000+ mention running part of the job in a spreadsheet or Excel.
Those details are the raw material of demand: a price someone already pays, a tool someone left, a workaround someone maintains. The 2026 synthetic-participant literature documents the opposite in simulated answers: generic praise, invented familiarity and “suspiciously precise balance between praise and criticism.”
“Terrible service and literally fraudulent billing, they're refusing to refund us thousand of dollars after an explicit cancellation request.” – Capterra review, 1 star
“NO support for the lowest tier plan ($35/month), they break parts of the product & refuse to tell me how to fix it, can't sync with multiple calendars, payments are slow, no option for overnight pet sitting in the clients home” – Capterra review, 2 stars
“Not much features included. Maintaining shared excel which is costing zero is same as what bigin provides” – G2 review
That last one is a willingness-to-pay data point no persona gives you: the buyer’s real alternative costs zero. Our report on industries still running on spreadsheets is built on exactly this signal.
Fairness matters here. Real feedback has its own bias. Only 8.5% of Capterra reviews are three stars or lower, which means review platforms over-represent happy customers, and the people who write one-star reviews are the most motivated to write at all. NN/g makes the mirror-image point about online data: people rarely review a gas station unless something went wrong.
The difference is that real bias is bias about real events. You can correct for it by reading across ratings, counting recurrence instead of trusting one loud post, and weighting by source. A simulated answer has no events underneath to correct toward. We explain how to weigh sources in multi-signal startup idea validation and how to validate a niche’s viability.
The academic evidence converged on the same answer this year. A systematic review of 182 studies on LLM-generated participants, from UXtweak Research and the Slovak University of Technology and summarized by The Voice of User, found low variability, stereotyping, positivity bias and hallucinated needs across every model and prompting technique it covered.
The same group then ran its own experiment, which is the most useful one for anyone tempted to test designs or messages on a synthetic panel. It is a preprint, not yet peer reviewed, so read it as strong evidence rather than settled science.
The May 2026 preprint Distorted Perspectives of LLM-Simulated Preferences took 29 real preference tests run by organizations on the UXtweak platform, with 2,073 real participants and 78 tasks, and asked an LLM steered to simulate those exact audiences to pick the preferred design.
| Measure | Result |
|---|---|
| Real preference tests replayed | 29 (2,073 participants, 78 tasks) |
| Tasks where AI and human distributions differed significantly | 44% |
| AI agreed with humans’ most popular option | 53% |
| Normalized entropy, AI vs humans (higher = flatter) | 0.93 vs 0.86 |
| With GPT-5.2 reasoning model: significant differences | 38% (not a significant improvement) |
| With GPT-5.2: first-choice agreement | 65% |
| Rich personas vs generic: divergence | 46% vs 44% |
| Simulating one individual at a time: divergence | 91% |
A 53% hit rate on picking the winner is the number most people quote. The more important result is how the model failed.
The simulations did not just guess wrong. They made the options look closer than they were. Human preferences had a normalized entropy of 0.86. The AI’s were 0.93, meaning flatter. When real people had a clear favourite, the model spread its vote and handed back a near tie.
For a founder, a tie is worse than a wrong answer. A wrong answer invites an argument. A tie quietly confirms whatever you already wanted to build, because nothing in the data says no. The Voice of User’s summary calls it laundering a decision. We would call it the most expensive kind of agreement, the kind you pay for later in churn.
This is the same failure that makes founders overbuild. Our analysis of how small your MVP should actually be shows what happens when nothing in the process says no.
The standard fix is to give the persona more detail: demographics, personality, screening answers. In the preprint, rich personas built from the real samples’ own screening data diverged from humans in 46% of tasks. Generic personas diverged in 44%. The difference was not significant.
Worse, simulating one detailed individual at a time pushed divergence to 91% of tasks, and the model became rigidly repetitive. The practical takeaway is simple. If a vendor’s pitch rests on how detailed its personas are, that detail is not where accuracy comes from.
“If a tool just says "pretend you're a 35-year-old marketing manager", then yes, that's mostly UI on top of ChatGPT. It only becomes meaningfully different when it adds grounded source data, keeps longitudinal state across sessions, and shows calibration against real user responses instead of vibes.” – r/AI_Agents
The strongest positive result in the field comes from Stanford-led work, LLM agents grounded in self-reports. Its authors built agents from two-hour interviews with a national sample of 1,052 Americans. On held-out General Social Survey items, interview-based agents reached 83% of participants’ own two-week test-retest consistency, survey-based agents 82%, and combined agents 86%. Demographics-only agents reached 74%.
That is impressive, and it proves the point. The agents got close to real people because they were built from two hours of real people. The more real data you feed a simulation, the better it gets, which means the valuable input was the real data all along. For a founder with no customers yet, the cheapest version of that real data already exists: what your future customers wrote about current tools. We cover how to find it in the best customer complaint databases.
Simulated users behave like attentive experts. In a MeasuringU tree test, ChatGPT-4 found the correct path for all ten tasks on at least one run and nine of ten on four runs, while 33 real participants averaged about 51% success. None of the humans got every task right.
The same study found something useful: ChatGPT’s predicted ease ratings correlated with the humans’ at r = .62. So a model can sometimes predict how hard people will think something is, while being useless at predicting whether they succeed. A researcher on r/UXResearch saw the same gap in their own side-by-side test:
“They behave as very attentive very smart user that notices all the small details on the page and reads everything, even small text. Humans are obviously not like that at all. We had comprehension checking questions, and where humans are 30% correct synthetic are 90% correct.” – r/UXResearch
Chat models are trained to be helpful and agreeable. When you ask a persona whether it would use your product, you are asking an agreeable system to evaluate an idea you clearly care about. NN/g tested exactly this with a concept for delivering medical samples by drone. The synthetic rep answered that it “would find this application very useful” and that it “could be a game-changer.”
In the same study, real learners admitted they rarely finished online courses and rarely used discussion forums. The synthetic learners claimed to finish everything and to love forums. This is the Mom Test problem at machine scale: even real people give polite answers about hypothetical products, and a model tuned for politeness is the politest respondent there is. We unpack the founder version in how to stop delusional thinking when validating.
“AIs don't buy things. AIs don't have weird use cases where they need to do performative dances and sacrifice things to executive whimsy. Real humans who pay for your products do.” – r/ProductManagement
“What people say, feel and do are all completely different things. These bots will be based entirely on what people say, no LLM can understand the nuance of tone, context and social biases.” – r/UXResearch
No. Demand is a behavior, not an opinion. It shows up as people complaining without being asked, paying for a workaround, switching tools, or handing over money. A synthetic user can only predict what someone might say if asked, and it predicts from the average of what has been written online.
That is why the evidence that matters most for founders is revealed, not stated. When a buyer writes that they are “maintaining shared excel which is costing zero,” they are telling you the real price of your competition. When 14,000+ reviews mention spreadsheets, that is a count of people doing the job by hand. Our guide to validating a startup idea ranks evidence by exactly this ladder, and finding problems to solve shows where to look for it. To size what you find, use how to calculate market size.
“But it's not evidence of anything. ChatGPT isn't your user. You're better off using the AI to crawl social media for secondary evidence from your target user group. At least then it's insight from people who could be your users, which is better than nothing.” – r/UXDesign
Not reliably. A persona never sees a card form. It will name a plausible price because plausible prices are all over its training data, and it will never feel the difference between $29 and $49. Research vendors are careful here too: one of the ranking guides written by a synthetic-research company states outright that synthetic outputs do not establish “exact willingness to pay.”
Real price evidence looks different. 23.3% of 1-2 star Capterra reviews mention price or cost, against 17.3% of the rest. 700+ reviews say a product is “not worth” it, overpriced or too expensive. 160+ describe a specific price hike. Each comes with a reason, which is what you actually need to set a price.
“They raise the price every year and in return customer service has become more unhelpful. The Technical Support has become non-existent, expect hold times of over an hour and don't bother calling you back or returning your emails.” – Capterra review, 2 stars
“The cost is very high. Also, we found that our job requirements were not specialized enough for the algorithm to really make an impact.” – G2 review
For a method that starts from real prices, see how to price a micro SaaS and our study of AI SaaS pricing models, which uses revenue data rather than stated preference.
The honest base rate for “will people pay?” is sobering, and no persona will give it to you. Of 8,600+ revenue-verified startups on TrustMRR, 50.6% earned any money in the last 30 days. 1,200+, or 14.0%, cleared $1,000. The median earner made $199.50.
Every one of those products presumably passed someone’s gut check. A synthetic panel would have liked most of them. The market paid for about half. If you want to know what products like yours actually earn, look it up in revenue intelligence rather than asking a simulation to imagine it. Our analysis of how long it takes to grow a SaaS shows the full distribution, and the AI SaaS revenue reality check covers AI products specifically. The revenue intelligence guide explains the fields.
“Most founders don’t fail because they can’t build. They fail because they build before validating the math.” – r/microsaas
Synthetic users earn their place in preparation, not decisions. Researchers who dislike them still name a short list of legitimate jobs, and NN/g’s guidance lands in the same place: desk research, hypothesis generation and piloting research instruments.
“Okay for rough "hey I had this idea" phase of ideation. But in no way should be used more than that or replacing research with humans.” – r/UXResearch
“I'd use synthetic users as a sparring partner, not as evidence. they're great for stress-testing assumptions early ("would this flow break for a first-time user?") and spotting obvious friction before recruiting real users.” – r/UXDesign
The three jobs below are the ones where a simulation saves real time without pretending to be a customer.
A persona is a decent first reader. It catches jargon, unclear value propositions and headlines that assume context the reader does not have. It will not tell you which message converts, because the preprint above shows simulated preferences flatten toward a tie. Use it to remove obvious confusion, then let real traffic decide.
The strongest version grounds the copy test in real language. Pull the exact phrases customers use in complaints and check whether your headline speaks to them. That is how we recommend writing a landing page in how to get your first customer and how to get your first 100 SaaS users.
This is the best use we have seen. Run your interview script against a persona and you will find leading questions, double-barrelled questions and gaps before you burn a real call on them. One practitioner described a loop that keeps the persona honest:
“So interview synthetic users first quickly to identify additional blind spots/concerns/ideas to discuss with real people, have a chat with real people informed by that, feed the transcript back into the persona to improve it.” – r/UXResearch
“I used AI to better understand what might be the KPIs and OKRs of my stakeholders, and other things. This helps to come to the meetings much more prepared. Still need the meeting with the real people, but AI does help to refine the questions to ask” – r/ProductManagement
Pair this with our list of customer discovery questions and the startup idea validation checklist. To find the people to call, see how to find your first SaaS customers.
Because simulated users behave like attentive experts, they are a reasonable proxy for an expert review. MeasuringU’s conclusion was blunt and useful: if ChatGPT cannot find something in your navigation, you probably have a problem. The reverse does not hold. A persona finding it proves nothing about a distracted human.
“I rely on AI for evaluating for heuristics and adherence to our design system and principles. It’s pretty good for that and I see no reason to create the additional abstraction of a synthetic user.” – r/UXResearch
The practical answer is a table, not a verdict. Match the task to the evidence that can actually answer it.
| Task | Synthetic users? | Better evidence |
|---|---|---|
| Is the problem real and painful? | No | Recurring complaints, severity, churn flags |
| Will people pay, and how much? | No | Revenue of similar products, price complaints, pre-orders |
| Why do customers leave competitors? | No | 1-2 star reviews, switching mentions |
| Which design or message wins? | Weak (53% hit rate) | Live A/B test, real preference test |
| Is my interview guide any good? | Yes | Then run it on real buyers |
| What objections might I hear? | Yes, as a list | Real objections from calls and reviews |
| Is my copy confusing? | Yes, first pass | Real traffic and replies |
| Who should I recruit? | Partly | Where complaints cluster by role and industry |
| Obvious usability issues | Partly (expert proxy) | Five real users on a prototype |
Practitioner opinion is lopsided. In the r/UXResearch, r/ProductManagement and r/UXDesign threads we read, the top-voted replies reject synthetic users as evidence and accept them, at most, as a brainstorming aid.
“Synthetic Users are fake research for non-researchers. It's snake oil.” – r/UXResearch
“I’d treat synthetic users as a brainstorming tool, not a research participant. They might help generate hypotheses or stress test questions, but the moment we use them as a substitute for real people, we’re basically asking a model to predict what humans might say based on what humans have already said.” – r/UXResearch
“Why even bother doing the research then? For the false sense of confidence?” – r/UXResearch
“UXR here: Please do not do this. Who gets to hear the complaints about your product? A subreddit? Customer Service? Your sales team? Talk to them.” – r/ProductManagement
“The most transformative insights I’ve gathered haven’t been what someone said, but what I observed in context of use. This is something AI cannot do and all the more reason research should be supported.” – r/UXDesign
The defenders make a narrower claim, and it is fair:
“A synthetic pass costs basically nothing and takes an hour, so you can run it on stuff that'd never get budget for real testing early explorations, throwaway variants, the stuff that dies in a Slack thread instead of getting evidence either way.” – r/UXDesign
“Products are far more successful when everyone is in alignment on what to build. It might be off the mark because of the lack of accuracy, but it will at least be cohesive.” – r/UXDesign
Alignment and speed are real benefits. They are not evidence of demand.
Less than the noise suggests. We searched every BigIdeasDB corpus for companies selling synthetic audiences, then hand-checked each keyword match because “AI persona” also matches chatbot characters and AI assistants.
| Corpus | Size | Synthetic audience tools | Real feedback, survey or research tools |
|---|---|---|---|
| Revenue-verified startups (TrustMRR) | 8,600+ | 2 (last-30-day revenue $79 and $0) | 100+ (45 earning, median $96, 8 above $1K) |
| Stripe Index | 30,000+ | 2 of 45 keyword matches | 430+ keyword matches |
| Funded companies | 17,000+ | 2 of 8 keyword matches, both founded 2025 | 48 keyword matches |
| AI connectors (ChatGPT and Claude) | 7,000+ | 1 | 180+ mention surveys or feedback |
Real-feedback tooling outnumbers synthetic tooling by roughly an order of magnitude in every corpus. Search interest spiked in 2026, but the companies that earn visible revenue are still the ones that help teams hear from real people. For how to read these directories yourself, see the Stripe Index tools and the funded companies tools. Crowding matters less than density, as our SaaS market saturation analysis shows.
Freelance demand points the same way. Among 5,300+ Upwork postings we track, at least 8 pay for real human respondents or for recruiters to find them, including recruiters for survey participants in Japan, the UAE, Singapore and Saudi Arabia, and paid video-diary studies. We found one posting about structuring persona data.
The sample is small, so read it as direction, not size. The direction is that teams with budgets still buy access to real people when the answer matters. Our state of freelance demand report covers the wider market, and validating SaaS demand with Upwork jobs shows how to use postings as a signal. The Upwork analysis guide covers the tool.
Synthetic users win on speed. That is not in dispute. The question is what you get for the time.
| Method | Time to first answer | Direct cost | Evidence type | Answers demand? |
|---|---|---|---|---|
| Synthetic persona in a chat model | Minutes | Near zero | Predicted opinion | No |
| Dedicated synthetic panel | Minutes to hours | Subscription | Predicted opinion at scale | No |
| Complaint and review mining | Hours | Low | Revealed behavior, unprompted | Partly, strongly on pain |
| Revenue benchmarks of similar products | Minutes | Low | Revealed spending | Yes, on category WTP |
| Five customer interviews | 1-2 weeks | Incentives and time | Stated, with context | Partly |
| Pre-order or paid pilot | 2-4 weeks | Landing page and outreach | Revealed commitment | Yes |
The cheap, fast option that also carries behavior is complaint mining. It sits between the synthetic panel and the interview on cost, and above both on evidence of pain. For more tools in this lane, see our list of the best tools to find customer pain points, market research tools for startups, the best idea validation tools and Reddit research tools for founders.
Here is the order we recommend. It uses AI where it is strong and real evidence where decisions get made.
If you are validating in a market you do not know, read how to validate a business idea in an industry you don’t know. For the full evidence ladder, our idea validation hub collects every method, and the SaaS idea validation tool guide shows how to run steps one and two in one sitting.
The honest reason founders reach for synthetic users is that recruiting is slow and awkward. But real customers have already written millions of words about their problems, without being asked, often angrily, on Reddit, Capterra, G2 and app stores. That is customer research with the recruiting already done.
Four places to start:
One researcher put the logic in a sentence: go to whoever already hears the complaints. For a founder without customers, that is the review corpus of the tools your customers use today. Our guide to turning G2 reviews into SaaS ideas walks through it, and validating a SaaS idea with real reviews turns it into a checklist.
“Many teams skip clinician interviews and go straight to prototyping. This causes a lot of later revisions and wasted time.” – r/medtech
The research is consistent on one point: simulations get better only when fed real, context-specific data. The review of 182 studies found the best-aligned approaches all relied on it. So if you are going to use a persona, give it the evidence first.
A simple pattern: export 20 to 50 real complaints for your category, paste them into Claude, Gemini or ChatGPT, and instruct the model to answer only from those quotes and to say “not in the evidence” when it cannot. You now have a summarizer of real customers, not a simulation of imagined ones. The BigIdeasDB MCP server does this directly inside your assistant, and the MCP setup guide takes about five minutes.
Our AI prompts for business ideas include evidence-first prompts you can reuse, and our LLM benchmark for pain-point extraction shows which models summarize complaints most faithfully. Prefer a chat grounded in data? The AI research chat answers from the database, not from memory.
Synthetic answers have fingerprints, and they matter even if you never buy a synthetic panel, because bots and AI-assisted respondents now pollute real surveys too. The preprint authors list the tells. We added two from our own complaint data.
“Replace the term synthetic users with AI slop and answer your own question.” – r/UXDesign
Take a real founder question from our Reddit pipeline:
“I’m validating a simple, web-only booking platform for hair and beauty salons in Canada. The angle is no commissions, flat monthly fee (CA$25–40), and French-first for Quebec, English supported.” – r/canadasmallbusiness
Ask a synthetic salon owner about it and you will hear that flat pricing is appealing and French support is valuable. Plausible, agreeable, and uncheckable. Now look at what 380+ Capterra reviews that mention salons, barbers or stylists actually say. Support comes up in 27.7% of them and price in 24.5%. 14.4% mention no-shows, reminders or text messages. Commission comes up in 1.6%.
“Only one thing is having to pay additionally to send sms to clients, It would be nice to have it included in the monthly price” – Capterra review, 5 stars
“The price tiers are not very inclusive. I would have to upgrade to a much higher tier to be able to take deposits from customers.” – Capterra review, 5 stars
Those are the details that change the plan. The founder’s angle was commissions, which barely appear. The real pricing pain is SMS reminders sold as an add-on and deposits locked behind higher tiers, both tied to no-shows. A flat fee that includes reminders and deposits is a sharper wedge than “no commissions”, and no persona would have surfaced it, because it comes from people who have paid for the add-on. The same method works for any vertical. Our small business software pain points report is a good place to start.
“No I wouldn't, because that would destroy my credibility and reputation.” – r/UXDesign
“If your problem is to learn what problems our finance team customers face then your approach would not be an option.” – r/ProductManagement
For more failure patterns, read our startup failure statistics, lessons from failed business ideas and the growth levers founders never pull.
All first-party numbers come from read-only SQL over BigIdeasDB’s live Supabase tables on September 25, 2026. Corpus counts are rounded down with a trailing plus.
| Source | Size | Used for | Limitation |
|---|---|---|---|
| Capterra reviews | 270,000+ | Theme shares by rating, specificity markers | Keyword tagging misses paraphrase and double-counts overlaps; only 8.5% are 3 stars or lower |
| Capterra pain points | 39,000+ | Churn flags and severity by category | AI-extracted labels; categories overlap (support vs service) |
| App Store and Google Play reviews | 136,000+ | Consumer theme shares | Sample skews to low ratings by design; app mix is not the whole store |
| G2 reviews and insights | 9,400+ insights | Quotes on price and support | Smaller corpus; structured Q&A format |
| Reddit pain-point pipeline | 2,300+ records | Founder quotes | Covers tracked subreddits only |
| TrustMRR revenue-verified startups | 8,600+ | Revenue base rates, tool revenue | Self-listed startups; skews to indie and small products |
| Stripe Index | 30,000+ | Synthetic vs feedback tool counts | Keyword-bounded directory sweep, not a census |
| Funded companies | 17,000+ | Funded synthetic-research count | Descriptions vary in detail; small counts |
| AI connector census | 7,000+ | Connector counts | Descriptions only; no usage data |
| Upwork postings | 5,300+ | Paid human-respondent demand | Small sample; titles only; no budget data used |
| arXiv 2605.18311 | 29 tests, 2,073 people | Preference-test accuracy | Preprint; GPT models only; one platform |
| arXiv 2411.10109 | 1,052 people | Best-case grounded agents | Survey items, not purchase behavior |
| MeasuringU tree test | 33 people | Too-smart finding | One tree, small sample |
| NN/g evaluation | 3 studies replayed | Sycophancy examples | 2024 models; qualitative |
| Google Trends | Relative index | Interest timing | Relative, not volume; worldwide |
Three things this page does not prove. First, we did not run our own synthetic panel against our review data, so the claim that personas cannot produce post-purchase complaints is structural (they have no purchase history) rather than measured on a specific tool. Second, keyword theme tagging is approximate: a review saying “great support, terrible billing” counts for both. Third, negative reviews are not a random sample of customers, which is exactly why we read them for what goes wrong rather than for how common it is overall.
We also could not verify one widely shared view count for a synthetic-research video, so we left it out. If a better study lands, we will update the numbers and the date.
BigIdeasDB holds 1M+ real complaints from Reddit, Capterra, G2 and app stores, plus revenue-verified data on 8,600+ startups and 30,000+ companies from Stripe’s directory. Filter to your category and read what customers already said, unprompted.
See BigIdeasDB plans →We built BigIdeasDB for the question synthetic users cannot answer: is this problem real, and will anyone pay? If you are doing customer research on a budget, this is the order we would use the tools in:
| Rank | Tool | Best for |
|---|---|---|
| 1 | BigIdeasDB | Real complaints, switching reasons, revenue benchmarks and market density in one place |
| 2 | ChatGPT | Drafting and piloting interview guides |
| 3 | Claude | Summarizing real complaints you paste in |
| 4 | Perplexity | Finding public threads and reviews to read |
| 5 | Notion | Keeping an interview and evidence log |
To go further: search complaints in the pain points database, learn how to use the pain points database, check an idea with the idea evaluator (the idea evaluator guide explains the scores), or find problems worth solving with our guide to finding problems worth solving. For the adjacent decisions, read how founders research markets, how to find product-market fit, competitive landscape analysis, and whether AI can write a business plan, plus competitor research tools and how to do competitor analysis. New to the platform? Start with what BigIdeasDB is.
Synthetic users are AI-generated personas that answer research questions as if they were a member of your target audience. You describe a segment, a large language model role-plays it, and you get interview transcripts or survey answers in minutes. They produce text about customers, not evidence from customers.
No. They can speed up preparation, but they cannot tell you whether a problem is painful enough to pay for. In a May 2026 preprint that replayed 29 real preference tests (2,073 people), LLM simulations picked the same favourite option as humans only 53% of the time, and the preference distributions differed significantly in 44% of tasks.
They are accurate on average opinions that are widely written about online and weak on behavior, priorities and niche audiences. The best published result, agents built from two-hour interviews with 1,052 Americans, reached 83% to 86% of people's own test-retest consistency on survey items. That required interviewing the real people first.
Because chat models are tuned to be helpful and agreeable, a trait researchers call sycophancy. NN/g found a synthetic user called a drone sample-delivery concept a potential game-changer, and synthetic learners claimed to love course forums that real learners said they never used. An AI persona has no budget, no switching cost and no reason to say no.
No. Demand is revealed by behavior: people complaining unprompted, paying for workarounds, switching tools and spending money. A persona can only predict what a customer might say. As of September 2026, 50.6% of 8,600+ revenue-verified startups earned any money in the last 30 days, which is the gap between plausible and paid.
Not reliably. A persona never faces a real card prompt. Real price evidence comes from what similar products earn and what buyers write when a price hurts. In 1-2 star Capterra reviews, 4.4% name a specific dollar amount, against 0.3% of other reviews, and those numbers come attached to a reason.
Preparation work: drafting and piloting interview guides, generating objections to rehearse, spotting confusing copy, brainstorming segments to recruit, and a first pass on obvious usability issues. Treat every output as a hypothesis to test with real people or real complaint data.
Only if you know what it cannot tell you. NN/g's position is that some of what you learn will be wrong and, without real research, you will not catch it. For founders there is a cheaper alternative to both: mining what real customers already wrote in reviews and forums, which costs little and carries real behavior.
Post-purchase experience. Across 9,400+ one and two star Capterra reviews, 62.1% mention support, billing, cancellation, bugs or updates, and only 9.9% mention a missing feature. In 91,000+ one and two star App Store reviews the split is 37.5% against 4.1%. A persona has never waited for a support ticket.
In the 2026 preference-test preprint, no. Detailed personas built from real screening data diverged from humans in 46% of tasks against 44% for generic ones, which was not a significant difference. Simulating one specific individual at a time made it worse, at 91% of tasks.
Not so far. Swapping GPT-4.1 for the GPT-5.2 reasoning model moved significant differences from 44% to 38% of tasks, which the authors did not find statistically significant. Lower temperature settings did not fix the flattened preferences either.
A proto-persona is a written assumption about who your customer is, used to align a team before research. A synthetic user is an interactive AI version you can question. Both are hypotheses. The proto-persona is at least honest about it, because nobody mistakes a one-page document for an interview.
Start with what customers already wrote. Filter reviews and forum posts to your category, count which complaints recur, read the one and two star reviews for switching reasons and price pain, then check what similar products earn. BigIdeasDB puts 1M+ complaint records and revenue data in one place for that.
Only if you label them clearly as simulated. Presenting synthetic output as customer research is misleading, and practitioners on r/UXDesign warn it destroys credibility once someone asks who you talked to. Real quotes and real numbers carry the room.
No. Synthetic data is generated records used to test software, train models or protect privacy, and it can be very useful. Synthetic users are simulated opinions used as a stand-in for customers. The first tests your system. The second claims to test your market.
It spiked and is cooling. Google Trends interest in "synthetic research" peaked the week of May 31, 2026 and now sits at 24% of that peak. "Synthetic users" peaked on May 17, 2026 and is at 30%. All three terms we checked stepped up sharply in August 2025.
BigIdeasDB is our pick for founders because it answers the demand question with real records: 1M+ complaints across Reddit, Capterra, G2 and app stores, plus revenue-verified startup data and a directory of 30,000+ companies. Generalist AI tools like ChatGPT and Claude are best used to summarize that evidence, not to invent it.
BigIdeasDB Research. (2026). Synthetic Users vs Real Customer Research: What AI Personas Can't Tell You. BigIdeasDB. Retrieved from https://bigideasdb.com/synthetic-users-vs-real-customer-research