Customer Research

Synthetic Users vs Real Customer Research: What AI Personas Can't Tell You

AI personas answer the question you ask before launch. Real customers leave over what happens after it. We measured the gap across 400,000+ reviews and the 2026 research.

28 min readShare →
62.1%
1-2 star Capterra reviews about post-purchase pain
9.9%
Same reviews citing a missing feature
53%
AI picked the humans' favourite design (2026 preprint)
50.6%
Revenue-verified startups that earned anything last month

Synthetic users can help you prepare for customer research. They cannot replace it. An AI persona answers the question you ask before launch: would you use this? Real customers leave over what happens after they pay. Across 9,400+ one and two star Capterra reviews, 62.1% are about support, billing, cancellation, bugs or updates. Only 9.9% name a missing feature.

That gap is the whole argument. A synthetic user has never waited three days for a support reply, never been charged after cancelling, never had an update break a workflow. So it cannot warn you about the things that actually make customers churn, and it cannot tell you whether anyone will pay. We checked the pattern against 91,000+ one and two star app reviews, the churn flags on 39,000+ documented pain points, and the 2026 academic research on synthetic participants.

This page is for founders and product people deciding whether AI personas can stand in for talking to customers. If you want the broader method of using AI for research, read how to use AI for market research. If you are about to get on calls, start with our customer discovery questions.

Key takeaways
  • In 9,400+ one and two star Capterra reviews, 62.1% cite post-purchase pain (support 42.7%, billing or cancelling 18.4%, bugs 17.8%, updates 8.9%). Only 9.9% cite a missing feature.
  • The same split holds in 91,000+ one and two star app reviews: 37.5% post-purchase against 4.1% missing features.
  • Documented pain points about customer service carry a churn flag 98.2% of the time, against 57.4% for feature limitations.
  • In a May 2026 preprint replaying 29 real preference tests, AI simulations matched the humans’ top choice 53% of the time. Richer personas did not help.
  • Only 50.6% of 8,600+ revenue-verified startups earned anything in the last 30 days. That, not a persona’s enthusiasm, is the base rate for demand.

Can synthetic users replace real customer research? The short answer

The short answer
No. Use synthetic users to prepare: pilot interview guides, rehearse objections, tighten copy. Use real evidence to decide: complaints customers wrote without being asked, what similar products earn, and a handful of conversations with people who already pay for a workaround. As of September 2026, 62.1% of one and two star Capterra reviews are about experiences an AI persona has never had, and the best published synthetic panels still needed real interviews to get close to human answers.

The rest of this page shows the numbers behind that answer, where synthetic users genuinely earn their place, and a workflow that gets you real customer voice without weeks of recruiting.

What are synthetic users, and what are they not?

A synthetic user is an AI-generated persona you can interview. You describe a segment, such as “operations managers at 20-person logistics firms”, and a large language model answers as that person. Dedicated tools wrap this in panels, dashboards and survey exports. A general chat model like ChatGPT or Claude does the same with a role prompt.

Three things get confused under the same label, and the difference matters:

  • Synthetic users or respondents: simulated opinions standing in for customers. This page is about these.
  • Proto-personas: written assumptions a team agrees on before research. Useful for alignment, honest about being guesses.
  • Synthetic data: generated records for testing software or training models. Often genuinely useful, and not a claim about your market.

The Nielsen Norman Group put the core problem bluntly in its evaluation of synthetic users: they provide “artificial research findings produced without studying real users.” The output looks like research. It is a prediction of what research might say. For a definition of who you should be studying in the first place, see our guide to the ideal customer profile for startups.

Why is everyone asking about synthetic users now?

Because the pitch got louder in 2026. On Google Trends, worldwide interest in “synthetic research” was flat for four years, stepped up in August 2025 and peaked the week of May 31, 2026. It now sits at 24% of that peak. “Synthetic users” peaked on May 17, 2026 and is at 30%. “AI personas” peaked the same week as synthetic research and is at 26%.

The most-watched YouTube talk on the subject, a July 2026 conference session on persona engineering, has 25,000+ views. NN/g’s explainer video has 12,000+. Four of the nine pages ranking for “synthetic users vs real users” are written by companies selling synthetic research, and most of the rest are UX-only. So most of what a founder reads on this topic comes from the people with the most to gain from a yes.

The founder version of the question is older than the tools. People have always wanted a shortcut past the awkward part:

“I want to test out my idea before putting money into it, but traditional market research takes too long and costs too much.” – r/solopreneurs
“How do we distinguish must-haves from nice-to-haves early?” – r/leanstartup

Synthetic users promise to answer both in an afternoon. The data below shows which half of that promise holds. If you are still choosing what to build, our guides to finding startup ideas and choosing between startup ideas come first.

How we measured the gap

Most comparisons of synthetic and real users test the simulation against a study. We took the other side: what do real customers actually talk about when a product fails them, and could a simulated customer have said it?

We ran read-only SQL over BigIdeasDB’s live tables on September 25, 2026. We split 270,000+ Capterra reviews and 136,000+ App Store and Google Play reviews by star rating, then tagged the text for themes with keyword patterns: support, billing and cancellation, bugs and outages, updates, price, missing features, integrations and ease of use. A review can match several themes. We also pulled churn flags from 39,000+ documented pain points, revenue base rates from 8,600+ revenue-verified startups on TrustMRR, and counted synthetic-research tools across the Stripe Index, funded companies, and the AI connector census.

Keyword tagging undercounts and overcounts at the edges. The full method and every limitation are in the methodology and data sources sections.

The headline finding: customers leave over things personas never lived

Real customers mostly complain about the relationship after purchase, not the product idea before it. In one and two star Capterra reviews, 62.1% mention support, billing, cancellation, bugs or updates. 36.6% mention only those themes and nothing about price, features, integrations or ease of use. Missing features show up in 9.9%.

A synthetic user interviewed about a concept will happily discuss features, workflows and pricing tiers, because that is what a concept interview is about. It will not bring up the support queue, the auto-renewal or the update that moved a button, because it has never used your product. The complaints that end customer relationships are the ones a simulation is structurally unable to produce.

“Everything. I have gotten zero help. Cannot login and cannot get anyone to answer my emails or call me back. I have been trying to cancel for months because no one will help me and I cannot get help to do that either! Very dissatisfied!” – Capterra review, 1 star

Nobody predicts that review in a concept test. It is still the review that decides whether the next buyer signs up, and it is the kind of evidence our business pain points research is built from.

What 9,400+ one and two star Capterra reviews are about

Support is the single largest theme in negative B2B software reviews, at 42.7%. Price is 18.6%. Missing features are 9.9%. Here is the full split.

ThemeShare of 1-2 star reviewsCan a synthetic user report it?
Support or customer service42.7%No. Requires having needed help
Price or cost18.6%Partly. Can guess, never pays
Billing, contracts, cancellation, refunds18.4%No. Requires a billing history
Bugs, outages, broken features17.8%No. Requires using the product
Missing features9.9%Yes, and it over-produces them
Integrations and sync9.9%Partly. Generic, not your stack
Updates or version changes8.9%No. Requires a before and after
Ease of use8.4%Partly. Opinion, not observed struggle
Any post-purchase theme62.1%No
Only post-purchase themes36.6%No
Share of 1-2 star Capterra reviews matching each theme (a review can match several). Source: BigIdeasDB Capterra corpus, 9,400+ low-rated reviews out of 270,000+. Queried September 25, 2026.

The feature row is the one synthetic research is good at, and it is the smallest. NN/g saw the same thing from the other direction: asked what makes a course engaging, its synthetic user listed seven factors with no sense of which mattered. “Real people care about some things more than others. Synthetic users seem to care about everything,” the team wrote. To see how that priority shows up in software categories, read our state of SaaS pain points report.

“Requests for help. I submitted several tickets for help and got no response. There is no way to cancel your account by yourself. You have to submit a help ticket for this, and of course, I got no response.” – Capterra review, 1 star
“Stability issues, Support issues, Consistency issues, Accuracy issues. We spent more time working on issues than using the product.” – Capterra review, 2 stars

The same pattern in 91,000+ one and two star app reviews

Consumer apps show the same shape at a different scale. In one and two star App Store and Google Play reviews, 37.5% mention billing, trials, bugs, updates, support or lost data. Missing features appear in 4.1% and ease-of-use complaints in 1.4%. That is 9 post-purchase complaints for every feature request.

Theme1-2 star reviews3-5 star reviews
Bugs, crashes, not loading14.4%n/a
Billing, trials, subscriptions, refunds14.2%n/a
Updates or new versions9.5%8.8%
Support8.0%n/a
Price, paywall, premium7.0%n/a
Ads6.7%6.1%
Missing features4.1%n/a
Leaving: cancel, uninstall, switch11.8%3.3%
Names a dollar amount5.0%2.0%
Share of 1-2 star app reviews matching each theme. Source: BigIdeasDB App Store and Google Play corpus, 91,000+ low-rated reviews out of 136,000+. Queried September 25, 2026.

Low-rated app reviews are 3.6x as likely to describe leaving as the rest. They are 2.5x as likely to name a dollar figure. That is behavior, recorded after the fact, with a number attached. For the full mobile picture, see our state of mobile app pain points.

“I downloaded the app to test out the free version and was immediately charged for a subscription upon opening the app for the first time. Now trying to figure out how to get a refund” – App Store review, 1 star
“I had a subscription for radarbox, after the update my app stopped working, had to reinstall. it does not recognize my subscription. asking to pay double for a new subscription. Basically scam.” – App Store review, 1 star

Support is the biggest complaint no persona will raise

Even on a narrow match (just “support” or “customer service”), 39.1% of 1-2 star Capterra reviews raise it, against 21.7% of higher ratings. The theme almost doubles as ratings fall. It is also the theme a concept interview never touches, because nobody needs support for a product they have not bought.

“Their support is dreadful. A week or two can go by with several emails to them and all we hear are crickets. There is no response sometimes until the third or fourth email. This is terribly unprofessional and is horrible for our agency and clients.” – Capterra review, 1 star
“No Support Company seems unstable Never get to speak to the same person from support twice Seems to be forgotten after purchase” – Capterra review, 2 stars
“Poor data. Influencers often miscategorized. No response from customer support.” – G2 review

“Forgotten after purchase” is a sentence no simulated buyer writes. For a founder, it is also a wedge: in crowded categories, a better support experience is often a stronger differentiator than another feature. Our customer support software limitations report shows how often that gap recurs, and what a SaaS moat looks like in the AI era explains why service is one of the few that survive.

Billing, trials and cancellation

18.4% of 1-2 star Capterra reviews and 14.2% of 1-2 star app reviews are about how money moved: auto-renewals, charges after cancelling, trials that billed early, refunds refused. A synthetic persona can opine on a price. It cannot be surprised by a charge.

“Information was not complete and they auto renewed me without any warning and we I canceled the same day they told me to pound sand.” – Capterra review, 1 star
“They said nothing about needing a plan, then charged me for a full year membership after I selected just the monthly plan. They would not provide a refund or change the plan length.” – Capterra review, 1 star
“I haven’t used or updated the app in nearly a year, and today, without any notification or heads up of any sort, they have charged me $119.99. They are refusing to refund. This is definitely fraudulent and I will be reporting this app.” – App Store review, 1 star
“It's hard to cancel and they charge weekly. After you send the cancelation request they still take money from your acct weekly. Scam!!!” – App Store review, 1 star

If you are pricing a product, these are the design decisions that decide your reviews. Our guide to what micro SaaS actually charges covers the price points. This data covers the part pricing pages leave out. Our app store database guide shows how to pull these billing complaints for any app category.

Updates that break things

8.9% of 1-2 star Capterra reviews and 9.5% of 1-2 star app reviews blame an update. This is the purest example of an experience that needs a before and an after. A synthetic user is always meeting your product for the first time.

“Since the recent update, everything feels messy and unnecessarily complicated. The interface is slower, harder to navigate, and many simple features are now buried or confusing to access. It’s made routine tasks more time-consuming and frustrating.” – Capterra review, 1 star
“The updates. Their timing is terrible. It is almost if they expect that you should have IT person on staff. They take no real responsibility for the actions that constantly cost their customers time and money.” – Capterra review, 2 stars
“Used to absolutely love this app. So much so that I paid for the premium version when it was a fixed rate. Now they’ve converted to a subscription and have gradually whittled down the features that I used to have.” – App Store review, 1 star

Which complaints actually predict churn?

Not all pain points are equal, and the ones that predict churn are post-purchase. Across 39,000+ documented Capterra pain points, those about customer service carry a churn-risk flag 98.2% of the time and customer support 96.7%. Feature limitations carry it 57.4% of the time.

Pain point categoryChurn-risk flagAvg severity
Customer service98.2%4.11
Customer support96.7%4.06
Pricing93.4%3.86
Integration issues87.2%4.01
Performance80.2%3.78
Reporting74.8%3.89
Functionality68.6%3.71
User experience67.5%3.75
Feature limitations57.4%3.65
Share of documented pain points flagged as a churn risk, with average severity (1-5), by category. Source: BigIdeasDB Capterra pain points, 39,000+ records. Queried September 25, 2026.

This is the priority signal synthetic panels flatten. A persona lists everything as important. Real evidence ranks it, and the ranking puts the relationship above the roadmap. If you are sorting your own backlog, our guide to prioritizing feature requests from customer reviews applies the same logic, and the Capterra analysis guide shows where these churn flags live.

Real complaints carry specifics synthetic answers invent

Angry reviews are specific. In 1-2 star Capterra reviews, 4.4% name a dollar amount, against 0.3% of other reviews, a 15x difference. 8.8% describe switching, migrating or cancelling, against 1.5%. Across all 270,000+ reviews, 14,000+ mention running part of the job in a spreadsheet or Excel.

Those details are the raw material of demand: a price someone already pays, a tool someone left, a workaround someone maintains. The 2026 synthetic-participant literature documents the opposite in simulated answers: generic praise, invented familiarity and “suspiciously precise balance between praise and criticism.”

“Terrible service and literally fraudulent billing, they're refusing to refund us thousand of dollars after an explicit cancellation request.” – Capterra review, 1 star
“NO support for the lowest tier plan ($35/month), they break parts of the product & refuse to tell me how to fix it, can't sync with multiple calendars, payments are slow, no option for overnight pet sitting in the clients home” – Capterra review, 2 stars
“Not much features included. Maintaining shared excel which is costing zero is same as what bigin provides” – G2 review

That last one is a willingness-to-pay data point no persona gives you: the buyer’s real alternative costs zero. Our report on industries still running on spreadsheets is built on exactly this signal.

Real reviews skew positive too

Fairness matters here. Real feedback has its own bias. Only 8.5% of Capterra reviews are three stars or lower, which means review platforms over-represent happy customers, and the people who write one-star reviews are the most motivated to write at all. NN/g makes the mirror-image point about online data: people rarely review a gas station unless something went wrong.

The difference is that real bias is bias about real events. You can correct for it by reading across ratings, counting recurrence instead of trusting one loud post, and weighting by source. A simulated answer has no events underneath to correct toward. We explain how to weigh sources in multi-signal startup idea validation and how to validate a niche’s viability.

What the 2026 research says about synthetic users

The academic evidence converged on the same answer this year. A systematic review of 182 studies on LLM-generated participants, from UXtweak Research and the Slovak University of Technology and summarized by The Voice of User, found low variability, stereotyping, positivity bias and hallucinated needs across every model and prompting technique it covered.

The same group then ran its own experiment, which is the most useful one for anyone tempted to test designs or messages on a synthetic panel. It is a preprint, not yet peer reviewed, so read it as strong evidence rather than settled science.

The preference-test experiment: right about as often as a coin flip

The May 2026 preprint Distorted Perspectives of LLM-Simulated Preferences took 29 real preference tests run by organizations on the UXtweak platform, with 2,073 real participants and 78 tasks, and asked an LLM steered to simulate those exact audiences to pick the preferred design.

MeasureResult
Real preference tests replayed29 (2,073 participants, 78 tasks)
Tasks where AI and human distributions differed significantly44%
AI agreed with humans’ most popular option53%
Normalized entropy, AI vs humans (higher = flatter)0.93 vs 0.86
With GPT-5.2 reasoning model: significant differences38% (not a significant improvement)
With GPT-5.2: first-choice agreement65%
Rich personas vs generic: divergence46% vs 44%
Simulating one individual at a time: divergence91%
Headline results from arXiv 2605.18311 (May 2026 preprint), as reported by the authors and summarized by The Voice of User. Not yet peer reviewed.

A 53% hit rate on picking the winner is the number most people quote. The more important result is how the model failed.

The false-tie problem

The simulations did not just guess wrong. They made the options look closer than they were. Human preferences had a normalized entropy of 0.86. The AI’s were 0.93, meaning flatter. When real people had a clear favourite, the model spread its vote and handed back a near tie.

For a founder, a tie is worse than a wrong answer. A wrong answer invites an argument. A tie quietly confirms whatever you already wanted to build, because nothing in the data says no. The Voice of User’s summary calls it laundering a decision. We would call it the most expensive kind of agreement, the kind you pay for later in churn.

This is the same failure that makes founders overbuild. Our analysis of how small your MVP should actually be shows what happens when nothing in the process says no.

Richer personas did not help

The standard fix is to give the persona more detail: demographics, personality, screening answers. In the preprint, rich personas built from the real samples’ own screening data diverged from humans in 46% of tasks. Generic personas diverged in 44%. The difference was not significant.

Worse, simulating one detailed individual at a time pushed divergence to 91% of tasks, and the model became rigidly repetitive. The practical takeaway is simple. If a vendor’s pitch rests on how detailed its personas are, that detail is not where accuracy comes from.

“If a tool just says "pretend you're a 35-year-old marketing manager", then yes, that's mostly UI on top of ChatGPT. It only becomes meaningfully different when it adds grounded source data, keeps longitudinal state across sessions, and shows calibration against real user responses instead of vibes.” – r/AI_Agents

The best case for synthetic users, and its catch

The strongest positive result in the field comes from Stanford-led work, LLM agents grounded in self-reports. Its authors built agents from two-hour interviews with a national sample of 1,052 Americans. On held-out General Social Survey items, interview-based agents reached 83% of participants’ own two-week test-retest consistency, survey-based agents 82%, and combined agents 86%. Demographics-only agents reached 74%.

That is impressive, and it proves the point. The agents got close to real people because they were built from two hours of real people. The more real data you feed a simulation, the better it gets, which means the valuable input was the real data all along. For a founder with no customers yet, the cheapest version of that real data already exists: what your future customers wrote about current tools. We cover how to find it in the best customer complaint databases.

Synthetic users are too smart to be your customers

Simulated users behave like attentive experts. In a MeasuringU tree test, ChatGPT-4 found the correct path for all ten tasks on at least one run and nine of ten on four runs, while 33 real participants averaged about 51% success. None of the humans got every task right.

The same study found something useful: ChatGPT’s predicted ease ratings correlated with the humans’ at r = .62. So a model can sometimes predict how hard people will think something is, while being useless at predicting whether they succeed. A researcher on r/UXResearch saw the same gap in their own side-by-side test:

“They behave as very attentive very smart user that notices all the small details on the page and reads everything, even small text. Humans are obviously not like that at all. We had comprehension checking questions, and where humans are 30% correct synthetic are 90% correct.” – r/UXResearch

Why an AI persona says yes to your idea

Chat models are trained to be helpful and agreeable. When you ask a persona whether it would use your product, you are asking an agreeable system to evaluate an idea you clearly care about. NN/g tested exactly this with a concept for delivering medical samples by drone. The synthetic rep answered that it “would find this application very useful” and that it “could be a game-changer.”

In the same study, real learners admitted they rarely finished online courses and rarely used discussion forums. The synthetic learners claimed to finish everything and to love forums. This is the Mom Test problem at machine scale: even real people give polite answers about hypothetical products, and a model tuned for politeness is the politest respondent there is. We unpack the founder version in how to stop delusional thinking when validating.

“AIs don't buy things. AIs don't have weird use cases where they need to do performative dances and sacrifice things to executive whimsy. Real humans who pay for your products do.” – r/ProductManagement
“What people say, feel and do are all completely different things. These bots will be based entirely on what people say, no LLM can understand the nuance of tone, context and social biases.” – r/UXResearch

Can synthetic users tell you whether there is demand?

No. Demand is a behavior, not an opinion. It shows up as people complaining without being asked, paying for a workaround, switching tools, or handing over money. A synthetic user can only predict what someone might say if asked, and it predicts from the average of what has been written online.

That is why the evidence that matters most for founders is revealed, not stated. When a buyer writes that they are “maintaining shared excel which is costing zero,” they are telling you the real price of your competition. When 14,000+ reviews mention spreadsheets, that is a count of people doing the job by hand. Our guide to validating a startup idea ranks evidence by exactly this ladder, and finding problems to solve shows where to look for it. To size what you find, use how to calculate market size.

“But it's not evidence of anything. ChatGPT isn't your user. You're better off using the AI to crawl social media for secondary evidence from your target user group. At least then it's insight from people who could be your users, which is better than nothing.” – r/UXDesign

Can synthetic users estimate willingness to pay?

Not reliably. A persona never sees a card form. It will name a plausible price because plausible prices are all over its training data, and it will never feel the difference between $29 and $49. Research vendors are careful here too: one of the ranking guides written by a synthetic-research company states outright that synthetic outputs do not establish “exact willingness to pay.”

Real price evidence looks different. 23.3% of 1-2 star Capterra reviews mention price or cost, against 17.3% of the rest. 700+ reviews say a product is “not worth” it, overpriced or too expensive. 160+ describe a specific price hike. Each comes with a reason, which is what you actually need to set a price.

“They raise the price every year and in return customer service has become more unhelpful. The Technical Support has become non-existent, expect hold times of over an hour and don't bother calling you back or returning your emails.” – Capterra review, 2 stars
“The cost is very high. Also, we found that our job requirements were not specialized enough for the algorithm to really make an impact.” – G2 review

For a method that starts from real prices, see how to price a micro SaaS and our study of AI SaaS pricing models, which uses revenue data rather than stated preference.

What real revenue data says instead

The honest base rate for “will people pay?” is sobering, and no persona will give it to you. Of 8,600+ revenue-verified startups on TrustMRR, 50.6% earned any money in the last 30 days. 1,200+, or 14.0%, cleared $1,000. The median earner made $199.50.

Every one of those products presumably passed someone’s gut check. A synthetic panel would have liked most of them. The market paid for about half. If you want to know what products like yours actually earn, look it up in revenue intelligence rather than asking a simulation to imagine it. Our analysis of how long it takes to grow a SaaS shows the full distribution, and the AI SaaS revenue reality check covers AI products specifically. The revenue intelligence guide explains the fields.

“Most founders don’t fail because they can’t build. They fail because they build before validating the math.” – r/microsaas

Where synthetic users are genuinely useful

Synthetic users earn their place in preparation, not decisions. Researchers who dislike them still name a short list of legitimate jobs, and NN/g’s guidance lands in the same place: desk research, hypothesis generation and piloting research instruments.

“Okay for rough "hey I had this idea" phase of ideation. But in no way should be used more than that or replacing research with humans.” – r/UXResearch
“I'd use synthetic users as a sparring partner, not as evidence. they're great for stress-testing assumptions early ("would this flow break for a first-time user?") and spotting obvious friction before recruiting real users.” – r/UXDesign

The three jobs below are the ones where a simulation saves real time without pretending to be a customer.

Message and copy testing

A persona is a decent first reader. It catches jargon, unclear value propositions and headlines that assume context the reader does not have. It will not tell you which message converts, because the preprint above shows simulated preferences flatten toward a tie. Use it to remove obvious confusion, then let real traffic decide.

The strongest version grounds the copy test in real language. Pull the exact phrases customers use in complaints and check whether your headline speaks to them. That is how we recommend writing a landing page in how to get your first customer and how to get your first 100 SaaS users.

Piloting interview guides

This is the best use we have seen. Run your interview script against a persona and you will find leading questions, double-barrelled questions and gaps before you burn a real call on them. One practitioner described a loop that keeps the persona honest:

“So interview synthetic users first quickly to identify additional blind spots/concerns/ideas to discuss with real people, have a chat with real people informed by that, feed the transcript back into the persona to improve it.” – r/UXResearch
“I used AI to better understand what might be the KPIs and OKRs of my stakeholders, and other things. This helps to come to the meetings much more prepared. Still need the meeting with the real people, but AI does help to refine the questions to ask” – r/ProductManagement

Pair this with our list of customer discovery questions and the startup idea validation checklist. To find the people to call, see how to find your first SaaS customers.

Heuristic flow checks

Because simulated users behave like attentive experts, they are a reasonable proxy for an expert review. MeasuringU’s conclusion was blunt and useful: if ChatGPT cannot find something in your navigation, you probably have a problem. The reverse does not hold. A persona finding it proves nothing about a distracted human.

“I rely on AI for evaluating for heuristics and adherence to our design system and principles. It’s pretty good for that and I see no reason to create the additional abstraction of a synthetic user.” – r/UXResearch

Which research task, which evidence?

The practical answer is a table, not a verdict. Match the task to the evidence that can actually answer it.

TaskSynthetic users?Better evidence
Is the problem real and painful?NoRecurring complaints, severity, churn flags
Will people pay, and how much?NoRevenue of similar products, price complaints, pre-orders
Why do customers leave competitors?No1-2 star reviews, switching mentions
Which design or message wins?Weak (53% hit rate)Live A/B test, real preference test
Is my interview guide any good?YesThen run it on real buyers
What objections might I hear?Yes, as a listReal objections from calls and reviews
Is my copy confusing?Yes, first passReal traffic and replies
Who should I recruit?PartlyWhere complaints cluster by role and industry
Obvious usability issuesPartly (expert proxy)Five real users on a prototype
Research tasks matched to the evidence that can answer them. Built from the BigIdeasDB findings and the 2024-2026 research cited on this page. September 2026.

What researchers and product managers say

Practitioner opinion is lopsided. In the r/UXResearch, r/ProductManagement and r/UXDesign threads we read, the top-voted replies reject synthetic users as evidence and accept them, at most, as a brainstorming aid.

“Synthetic Users are fake research for non-researchers. It's snake oil.” – r/UXResearch
“I’d treat synthetic users as a brainstorming tool, not a research participant. They might help generate hypotheses or stress test questions, but the moment we use them as a substitute for real people, we’re basically asking a model to predict what humans might say based on what humans have already said.” – r/UXResearch
“Why even bother doing the research then? For the false sense of confidence?” – r/UXResearch
“UXR here: Please do not do this. Who gets to hear the complaints about your product? A subreddit? Customer Service? Your sales team? Talk to them.” – r/ProductManagement
“The most transformative insights I’ve gathered haven’t been what someone said, but what I observed in context of use. This is something AI cannot do and all the more reason research should be supported.” – r/UXDesign

The defenders make a narrower claim, and it is fair:

“A synthetic pass costs basically nothing and takes an hour, so you can run it on stuff that'd never get budget for real testing early explorations, throwaway variants, the stuff that dies in a Slack thread instead of getting evidence either way.” – r/UXDesign
“Products are far more successful when everyone is in alignment on what to build. It might be off the mark because of the lack of accuracy, but it will at least be cohesive.” – r/UXDesign

Alignment and speed are real benefits. They are not evidence of demand.

Is anyone actually paying for synthetic research?

Less than the noise suggests. We searched every BigIdeasDB corpus for companies selling synthetic audiences, then hand-checked each keyword match because “AI persona” also matches chatbot characters and AI assistants.

CorpusSizeSynthetic audience toolsReal feedback, survey or research tools
Revenue-verified startups (TrustMRR)8,600+2 (last-30-day revenue $79 and $0)100+ (45 earning, median $96, 8 above $1K)
Stripe Index30,000+2 of 45 keyword matches430+ keyword matches
Funded companies17,000+2 of 8 keyword matches, both founded 202548 keyword matches
AI connectors (ChatGPT and Claude)7,000+1180+ mention surveys or feedback
Synthetic-research products vs real-feedback products across BigIdeasDB corpora. Synthetic counts are hand-reviewed keyword matches. Source: BigIdeasDB, queried September 25, 2026.

Real-feedback tooling outnumbers synthetic tooling by roughly an order of magnitude in every corpus. Search interest spiked in 2026, but the companies that earn visible revenue are still the ones that help teams hear from real people. For how to read these directories yourself, see the Stripe Index tools and the funded companies tools. Crowding matters less than density, as our SaaS market saturation analysis shows.

People still pay to recruit real humans

Freelance demand points the same way. Among 5,300+ Upwork postings we track, at least 8 pay for real human respondents or for recruiters to find them, including recruiters for survey participants in Japan, the UAE, Singapore and Saudi Arabia, and paid video-diary studies. We found one posting about structuring persona data.

The sample is small, so read it as direction, not size. The direction is that teams with budgets still buy access to real people when the answer matters. Our state of freelance demand report covers the wider market, and validating SaaS demand with Upwork jobs shows how to use postings as a signal. The Upwork analysis guide covers the tool.

Cost and speed compared

Synthetic users win on speed. That is not in dispute. The question is what you get for the time.

MethodTime to first answerDirect costEvidence typeAnswers demand?
Synthetic persona in a chat modelMinutesNear zeroPredicted opinionNo
Dedicated synthetic panelMinutes to hoursSubscriptionPredicted opinion at scaleNo
Complaint and review miningHoursLowRevealed behavior, unpromptedPartly, strongly on pain
Revenue benchmarks of similar productsMinutesLowRevealed spendingYes, on category WTP
Five customer interviews1-2 weeksIncentives and timeStated, with contextPartly
Pre-order or paid pilot2-4 weeksLanding page and outreachRevealed commitmentYes
Research methods compared by time, cost and the kind of evidence each produces. Qualitative comparison from the sources on this page. September 2026.

The cheap, fast option that also carries behavior is complaint mining. It sits between the synthetic panel and the interview on cost, and above both on evidence of pain. For more tools in this lane, see our list of the best tools to find customer pain points, market research tools for startups, the best idea validation tools and Reddit research tools for founders.

A hybrid workflow for founders

Here is the order we recommend. It uses AI where it is strong and real evidence where decisions get made.

  1. Mine real complaints first. Filter reviews and forum posts to your category in the pain points database. Count what recurs. Read the one and two star reviews for switching reasons.
  2. Check revenue reality. Look up what similar products earn in TrustMRR and how crowded the space is in the Stripe Index.
  3. Ground the persona. If you use an AI persona, paste 20 to 30 real complaint quotes into its context. It now argues from evidence, not from averages.
  4. Prepare with the persona. Pilot your interview guide, list objections, tighten your landing page copy.
  5. Talk to five real buyers. Find them where the complaints came from. Ask about the last time the problem happened and what they spent to fix it.
  6. Decide on behavior. Count pre-orders, paid pilots, switching and workarounds as demand. Treat everything simulated as a hypothesis.

If you are validating in a market you do not know, read how to validate a business idea in an industry you don’t know. For the full evidence ladder, our idea validation hub collects every method, and the SaaS idea validation tool guide shows how to run steps one and two in one sitting.

How to get real customer voice without recruiting anyone

The honest reason founders reach for synthetic users is that recruiting is slow and awkward. But real customers have already written millions of words about their problems, without being asked, often angrily, on Reddit, Capterra, G2 and app stores. That is customer research with the recruiting already done.

Four places to start:

One researcher put the logic in a sentence: go to whoever already hears the complaints. For a founder without customers, that is the review corpus of the tools your customers use today. Our guide to turning G2 reviews into SaaS ideas walks through it, and validating a SaaS idea with real reviews turns it into a checklist.

“Many teams skip clinician interviews and go straight to prototyping. This causes a lot of later revisions and wasted time.” – r/medtech

If you use an AI persona, ground it in real complaints

The research is consistent on one point: simulations get better only when fed real, context-specific data. The review of 182 studies found the best-aligned approaches all relied on it. So if you are going to use a persona, give it the evidence first.

A simple pattern: export 20 to 50 real complaints for your category, paste them into Claude, Gemini or ChatGPT, and instruct the model to answer only from those quotes and to say “not in the evidence” when it cannot. You now have a summarizer of real customers, not a simulation of imagined ones. The BigIdeasDB MCP server does this directly inside your assistant, and the MCP setup guide takes about five minutes.

Our AI prompts for business ideas include evidence-first prompts you can reuse, and our LLM benchmark for pain-point extraction shows which models summarize complaints most faithfully. Prefer a chat grounded in data? The AI research chat answers from the database, not from memory.

Red flags that your research is synthetic

Synthetic answers have fingerprints, and they matter even if you never buy a synthetic panel, because bots and AI-assisted respondents now pollute real surveys too. The preprint authors list the tells. We added two from our own complaint data.

  • Praise for things that are identical across every option.
  • Generic properties (“clarity”, “intuitive”) with no consequence drawn from them.
  • A suspiciously even balance of one positive and one negative.
  • Claimed familiarity (“like other sites I use”) with no specifics.
  • Guidelines instead of preferences, such as general accessibility advice.
  • No dollar amounts, no named tools, no workaround. Real angry reviews name a price 15x as often.
  • No support, billing or update story. In real negative reviews those appear 62.1% of the time.
“Replace the term synthetic users with AI slop and answer your own question.” – r/UXDesign

Worked example: a booking tool for salons

Take a real founder question from our Reddit pipeline:

“I’m validating a simple, web-only booking platform for hair and beauty salons in Canada. The angle is no commissions, flat monthly fee (CA$25–40), and French-first for Quebec, English supported.” – r/canadasmallbusiness

Ask a synthetic salon owner about it and you will hear that flat pricing is appealing and French support is valuable. Plausible, agreeable, and uncheckable. Now look at what 380+ Capterra reviews that mention salons, barbers or stylists actually say. Support comes up in 27.7% of them and price in 24.5%. 14.4% mention no-shows, reminders or text messages. Commission comes up in 1.6%.

“Only one thing is having to pay additionally to send sms to clients, It would be nice to have it included in the monthly price” – Capterra review, 5 stars
“The price tiers are not very inclusive. I would have to upgrade to a much higher tier to be able to take deposits from customers.” – Capterra review, 5 stars

Those are the details that change the plan. The founder’s angle was commissions, which barely appear. The real pricing pain is SMS reminders sold as an add-on and deposits locked behind higher tiers, both tied to no-shows. A flat fee that includes reminders and deposits is a sharper wedge than “no commissions”, and no persona would have surfaced it, because it comes from people who have paid for the add-on. The same method works for any vertical. Our small business software pain points report is a good place to start.

Mistakes to avoid

  • Counting synthetic responses as a sample. Ten thousand simulated answers are one model’s opinion repeated.
  • Testing your idea on a persona and calling it validation. Agreeable systems agree. See how to stop delusional thinking in validation.
  • Presenting simulated findings as customer research. Label them. Practitioners say unlabeled synthetic findings wreck credibility.
  • Using personas for niche audiences. NN/g warns synthetic data is patchiest where your users are specialized, which is where most B2B wedges live.
  • Skipping the post-purchase question. Ask real users about support, billing and updates in competitors. That is where 62.1% of angry reviews live.
  • Mistaking believability for accuracy. Experts could not tell synthetic personas from human ones on the surface. The problem is underneath.
“No I wouldn't, because that would destroy my credibility and reputation.” – r/UXDesign
“If your problem is to learn what problems our finance team customers face then your approach would not be an option.” – r/ProductManagement

For more failure patterns, read our startup failure statistics, lessons from failed business ideas and the growth levers founders never pull.

Methodology

All first-party numbers come from read-only SQL over BigIdeasDB’s live Supabase tables on September 25, 2026. Corpus counts are rounded down with a trailing plus.

  • Review themes. We concatenated each Capterra review’s cons, body and title, and each app review’s title and text, lower-cased them and matched regular expressions for each theme (for example support: “support”, “customer service”, “no response”, “ticket”). “Post-purchase” means support, billing or cancellation, bugs, updates or lost data. Shares are computed within rating bands.
  • Specificity markers. Dollar amounts match a dollar sign followed by a digit. Switching matches switched, migrating, moved to, left for or cancel.
  • Churn flags. Taken from the AI-extracted Capterra pain-point records, grouped by their category label.
  • Revenue base rates. Last-30-day revenue for all revenue-verified startups on TrustMRR.
  • Synthetic tool counts. Keyword matches for synthetic users, respondents, audiences, AI personas and simulated customers, then hand-reviewed; chatbots, companions and assistants were excluded.
  • External studies. Figures from arXiv 2605.18311 are as reported by its authors and The Voice of User; figures from arXiv 2411.10109 and MeasuringU are from the primary pages.

Data sources and their limits

SourceSizeUsed forLimitation
Capterra reviews270,000+Theme shares by rating, specificity markersKeyword tagging misses paraphrase and double-counts overlaps; only 8.5% are 3 stars or lower
Capterra pain points39,000+Churn flags and severity by categoryAI-extracted labels; categories overlap (support vs service)
App Store and Google Play reviews136,000+Consumer theme sharesSample skews to low ratings by design; app mix is not the whole store
G2 reviews and insights9,400+ insightsQuotes on price and supportSmaller corpus; structured Q&A format
Reddit pain-point pipeline2,300+ recordsFounder quotesCovers tracked subreddits only
TrustMRR revenue-verified startups8,600+Revenue base rates, tool revenueSelf-listed startups; skews to indie and small products
Stripe Index30,000+Synthetic vs feedback tool countsKeyword-bounded directory sweep, not a census
Funded companies17,000+Funded synthetic-research countDescriptions vary in detail; small counts
AI connector census7,000+Connector countsDescriptions only; no usage data
Upwork postings5,300+Paid human-respondent demandSmall sample; titles only; no budget data used
arXiv 2605.1831129 tests, 2,073 peoplePreference-test accuracyPreprint; GPT models only; one platform
arXiv 2411.101091,052 peopleBest-case grounded agentsSurvey items, not purchase behavior
MeasuringU tree test33 peopleToo-smart findingOne tree, small sample
NN/g evaluation3 studies replayedSycophancy examples2024 models; qualitative
Google TrendsRelative indexInterest timingRelative, not volume; worldwide
Every source used on this page, with its specific limitation. Snapshot September 25, 2026.

Coverage honesty

Three things this page does not prove. First, we did not run our own synthetic panel against our review data, so the claim that personas cannot produce post-purchase complaints is structural (they have no purchase history) rather than measured on a specific tool. Second, keyword theme tagging is approximate: a review saying “great support, terrible billing” counts for both. Third, negative reviews are not a random sample of customers, which is exactly why we read them for what goes wrong rather than for how common it is overall.

We also could not verify one widely shared view count for a synthetic-research video, so we left it out. If a better study lands, we will update the numbers and the date.

Hear from real customers in minutes, not weeks

BigIdeasDB holds 1M+ real complaints from Reddit, Capterra, G2 and app stores, plus revenue-verified data on 8,600+ startups and 30,000+ companies from Stripe’s directory. Filter to your category and read what customers already said, unprompted.

See BigIdeasDB plans →

Where BigIdeasDB fits

We built BigIdeasDB for the question synthetic users cannot answer: is this problem real, and will anyone pay? If you are doing customer research on a budget, this is the order we would use the tools in:

RankToolBest for
1BigIdeasDBReal complaints, switching reasons, revenue benchmarks and market density in one place
2ChatGPTDrafting and piloting interview guides
3ClaudeSummarizing real complaints you paste in
4PerplexityFinding public threads and reviews to read
5NotionKeeping an interview and evidence log
Tools for founder customer research, ranked by how much real customer evidence each gives you. September 2026.

To go further: search complaints in the pain points database, learn how to use the pain points database, check an idea with the idea evaluator (the idea evaluator guide explains the scores), or find problems worth solving with our guide to finding problems worth solving. For the adjacent decisions, read how founders research markets, how to find product-market fit, competitive landscape analysis, and whether AI can write a business plan, plus competitor research tools and how to do competitor analysis. New to the platform? Start with what BigIdeasDB is.

Frequently asked questions

What are synthetic users?

Synthetic users are AI-generated personas that answer research questions as if they were a member of your target audience. You describe a segment, a large language model role-plays it, and you get interview transcripts or survey answers in minutes. They produce text about customers, not evidence from customers.

Can synthetic users replace real customer research?

No. They can speed up preparation, but they cannot tell you whether a problem is painful enough to pay for. In a May 2026 preprint that replayed 29 real preference tests (2,073 people), LLM simulations picked the same favourite option as humans only 53% of the time, and the preference distributions differed significantly in 44% of tasks.

Are synthetic users accurate?

They are accurate on average opinions that are widely written about online and weak on behavior, priorities and niche audiences. The best published result, agents built from two-hour interviews with 1,052 Americans, reached 83% to 86% of people's own test-retest consistency on survey items. That required interviewing the real people first.

Why do AI personas agree with my startup idea?

Because chat models are tuned to be helpful and agreeable, a trait researchers call sycophancy. NN/g found a synthetic user called a drone sample-delivery concept a potential game-changer, and synthetic learners claimed to love course forums that real learners said they never used. An AI persona has no budget, no switching cost and no reason to say no.

Can synthetic users tell me if there is demand for my product?

No. Demand is revealed by behavior: people complaining unprompted, paying for workarounds, switching tools and spending money. A persona can only predict what a customer might say. As of September 2026, 50.6% of 8,600+ revenue-verified startups earned any money in the last 30 days, which is the gap between plausible and paid.

Can AI personas estimate willingness to pay?

Not reliably. A persona never faces a real card prompt. Real price evidence comes from what similar products earn and what buyers write when a price hurts. In 1-2 star Capterra reviews, 4.4% name a specific dollar amount, against 0.3% of other reviews, and those numbers come attached to a reason.

What are synthetic users actually good for?

Preparation work: drafting and piloting interview guides, generating objections to rehearse, spotting confusing copy, brainstorming segments to recruit, and a first pass on obvious usability issues. Treat every output as a hypothesis to test with real people or real complaint data.

Is synthetic research better than no research?

Only if you know what it cannot tell you. NN/g's position is that some of what you learn will be wrong and, without real research, you will not catch it. For founders there is a cheaper alternative to both: mining what real customers already wrote in reviews and forums, which costs little and carries real behavior.

What do real customers complain about that synthetic users miss?

Post-purchase experience. Across 9,400+ one and two star Capterra reviews, 62.1% mention support, billing, cancellation, bugs or updates, and only 9.9% mention a missing feature. In 91,000+ one and two star App Store reviews the split is 37.5% against 4.1%. A persona has never waited for a support ticket.

Do richer, more detailed personas make synthetic users more accurate?

In the 2026 preference-test preprint, no. Detailed personas built from real screening data diverged from humans in 46% of tasks against 44% for generic ones, which was not a significant difference. Simulating one specific individual at a time made it worse, at 91% of tasks.

Does a better model fix synthetic user accuracy?

Not so far. Swapping GPT-4.1 for the GPT-5.2 reasoning model moved significant differences from 44% to 38% of tasks, which the authors did not find statistically significant. Lower temperature settings did not fix the flattened preferences either.

What is the difference between a proto-persona and a synthetic user?

A proto-persona is a written assumption about who your customer is, used to align a team before research. A synthetic user is an interactive AI version you can question. Both are hypotheses. The proto-persona is at least honest about it, because nobody mistakes a one-page document for an interview.

How do I do customer research without recruiting participants?

Start with what customers already wrote. Filter reviews and forum posts to your category, count which complaints recur, read the one and two star reviews for switching reasons and price pain, then check what similar products earn. BigIdeasDB puts 1M+ complaint records and revenue data in one place for that.

Should I show synthetic user findings to stakeholders or investors?

Only if you label them clearly as simulated. Presenting synthetic output as customer research is misleading, and practitioners on r/UXDesign warn it destroys credibility once someone asks who you talked to. Real quotes and real numbers carry the room.

Are synthetic users the same as synthetic data?

No. Synthetic data is generated records used to test software, train models or protect privacy, and it can be very useful. Synthetic users are simulated opinions used as a stand-in for customers. The first tests your system. The second claims to test your market.

Is interest in synthetic research growing?

It spiked and is cooling. Google Trends interest in "synthetic research" peaked the week of May 31, 2026 and now sits at 24% of that peak. "Synthetic users" peaked on May 17, 2026 and is at 30%. All three terms we checked stepped up sharply in August 2025.

What is the best tool for real customer research in 2026?

BigIdeasDB is our pick for founders because it answers the demand question with real records: 1M+ complaints across Reddit, Capterra, G2 and app stores, plus revenue-verified startup data and a directory of 30,000+ companies. Generalist AI tools like ChatGPT and Claude are best used to summarize that evidence, not to invent it.

Cite this page
Last verified: September 25, 2026
BigIdeasDB Research. (2026). Synthetic Users vs Real Customer Research: What AI Personas Can't Tell You. BigIdeasDB. Retrieved from https://bigideasdb.com/synthetic-users-vs-real-customer-research
Founder, BigIdeasDB
Share →
Keep reading