Synthetic users get the average right and lose the difference between people
keywords: user research, synthetic users, applied ai, ux research
It's been keeping me up, watching research with real people get swapped for made-up users. Pull up a chair, this one's going to run long.

"I'm sure this difficulty is just mine, at my 63 years of age."
That's what a client wrote in the open-text field of a survey, after telling us that the trading screen — the home broker, as we call it in Brazil — frightened him, packed with information on every side. He apologised. For a problem that was mine.
The sentence is from 2019. I had joined a brokerage to lead the digital transformation of its services and platforms, and the redesign of the logged-in area, app and web, was part of that work. I still carry that sentence around. Yes, seven years have gone by.
Why was nobody transacting through the app?
Back then we had a number in hand and no explanation for it. The app was used to check things and almost never to transact: people logged in, looked at their balance, and left. Investing, redeeming, moving money, all of that happened on the web. An entire product live, underused, serving mostly to display a balance.
Hypotheses were not in short supply. The purchase flow was too long. The app was missing features. Investors are desktop creatures, spreadsheet open and chart alongside. A phone screen is too small to analyse an asset. All plausible. All easy to defend on a slide. And none of them explained the number.
The answer came from another client, in the same survey:
"I tried to add a fund through the app, it was confirmed in the app but when I logged into the website it wasn't there. I had to do it on the website. So the app doesn't give you a feeling of safety for even the smallest things. For now I only use it to check my balance."
It was all there. Nobody trusted that screen to move money. For looking, it was fine. And when the subject is investing, trust is the product you're selling. We were hunting for a flow problem and we had a credibility problem.
The closed-ended data from the same survey pointed the same way. We asked where people preferred to do each thing. Investing sat on the web. Checking a balance sat in the app. In the head-to-head, "faster" and "more practical" went to the app, and "safer" went to the website almost every time. The phone was more convenient and less trustworthy at once. Money went where the trust was.
And the gentleman in his sixties wasn't alone. Another client said he had trouble understanding the information on the portfolio screen and added: "that must be a shortcoming of mine."
Two people apologising for a problem that was ours.
That changed the project's priority. We stopped treating the app as a place to look things up and started treating it as a place to transact, which meant fixing trust before fixing features. It also changed how we talked to people: we went after the detractors, who in the Net Promoter Score are the ones scoring 0 to 6 on the 0-to-10 recommendation question, and who concentrate the churn risk. Talking to a detractor is uncomfortable, and it pays. Someone who is satisfied can't tell you what nearly went wrong. Someone on their way out can, and they tell you without diplomacy.
And what came of it?
This is the part I used to leave out, and shouldn't. The redesign went across every platform, and the app's rating went from 3.5 to 4.7 on both stores, App Store and Google Play. For a stretch we held a rating above the largest player in the market, something that looked out of reach when the project started. The target for new accounts opened was met afterwards too. And the part that matters most to me: people who used to log in only to look at a balance started investing from their phones. The behaviour we wanted to unlock did unlock.
Now the caveat, because this is where everyone overreaches. The redesign touched a lot of things at once, and other workstreams were running alongside it, including the one on account signup. I have no way of isolating how much of the rating and how much of those new accounts came from fixing trust, and how much came from performance, flow, content or marketing. What I'll defend is narrower: the project's priority changed because of those sentences, and the numbers moved along with it.
And here I get to what's been keeping me up.
What a synthetic user hands back
There's a promise going around of swapping research with real people for synthetic users, which means asking a language model to answer as if it were your customer. Fast, cheap, available at two in the morning. I use AI every day and I have no appetite for a speech about resisting it. But am I overreacting when I think we're throwing out precisely the part that makes the work worth something?
Look at what would have happened on that project. The plausible explanations on my list — long flow, missing features, desktop habit, small screen — are exactly the ones written all over the internet about investment apps. That's what a model is made of, and it would have handed me every one of them with confidence.
And I'll be honest to the end, because this objection is a good one: complaints about operations that don't sync and apps that don't feel safe are also written out there, in store reviews and on complaint sites. The point isn't that nobody ever said it. The point is that a model has no way of knowing which of those explanations is happening in your base, with those people, in that quarter. It hands you the list of possibilities. It doesn't point to which one is yours.
What research has already measured about synthetic answers
A much-cited study on this is the one by James Bisbee and colleagues, published in 2024 in Political Analysis. They asked ChatGPT to adopt personas and answer opinion questions, then compared the output with the American National Election Study, a real election survey run with people. The averages lined up nicely. Sounds great, and this is where the problem lives: the synthetic answers varied less than the real ones, regression coefficients often came out different, small changes in prompt wording shifted the distribution of answers, and the same prompt returned significantly different results across a three-month window. Fair to the study: the test ran in 2023, on GPT-3.5, and it isn't a portrait of today's models. The pattern still interests me a great deal: the tool got the aggregate averages right and lost the difference between people, including inside each subgroup, which is what those differing coefficients are telling us. Except my work lives off the difference between people.
The second study is harsher. Angelina Wang, Jamie Morgenstern and John Dickerson published a series of studies in 2025 in Nature Machine Intelligence, across four models and 3,200 real participants in total, covering sixteen demographic identities. The conclusion: models can misportray these groups and flatten the differences that exist within them. They tested mitigation techniques, which reduce the problem without removing it. And they urge caution whenever the identity of the participants matters to the task. For use as a supplement, a pilot study for instance, they even offer a path. The decision to replace, in their words, rests with whoever deploys it, weighing whether the benefits outweigh the harms.
Who disappears from the simulation
Translated into our day-to-day: people who don't write on the internet barely show up in the simulation, and when they do, they show up distorted. Older people, people on low incomes, anyone speaking a variety far from the standard, anyone with a disability. The way that 63-year-old said what he felt, apologising for his own difficulty, is the kind of speech that carries almost no weight in a model's training. And he was exactly the person I needed to hear, because that brokerage's active base had a lot of people his age.
How many research professionals use synthetic users today?
Few, and most are wary. In May 2026, User Interviews heard from 150 research professionals in an online survey, plus five in-depth interviews: 8% use AI tools to simulate participants, 21% have experimented, nearly half describe themselves as sceptical and 17% are outright opposed, while 80% say they use AI regularly in their work. Two honest caveats: the sample is a convenience one, recruited through the company's own channels, and that company makes its money recruiting human participants. Read it through that filter. Even so, the picture is of a profession that adopted AI for nearly everything and stalls at the point of replacing people.
Worth citing as well the State of UX 2026 from the Nielsen Norman Group, from January of this year, which doesn't touch synthetic users at all and argues the opposite: the more AI there is inside the product, the deeper you need to understand the people using it.
So where do synthetic users belong?
Now the other side, because criticism without an alternative is useless. Synthetic users have their place. Truly. Rehearsing a script before spending a real participant's hour. Generating hypotheses to test with people afterwards. Training whoever is going to moderate the conversation. Sweeping edge cases to work out what to ask. In those cases they genuinely save time.

The short version of the test I use to decide: ask who finds the error, and when. If the one who finds it is you, next week, in a conversation with a real person that's already on the calendar, use it without fear. If the one who finds it is your customer, on launch day, don't.
And there's something this test doesn't cover, so let me say it: it catches the wrong answer, not the question that never existed. A synthetic user answers what you asked well, and has no way of warning you about the subject that never crossed your mind. For that one I haven't found a shortcut. The tool prepares the fieldwork. The fieldwork it doesn't do.
The incentive is what genuinely worries me. Real people delay the schedule, disagree with the roadmap, and bring up a problem nobody wants to solve this quarter. A synthetic user answers in seconds, fits the sprint, and agrees with you. And that's no accident: a model trained to be helpful pulls towards agreement. NN/g tested synthetic users against their own studies with real people and recorded exactly that, the synthetic one appearing to care about everything put in front of it, which destroys any attempt at prioritisation. That's when research stops producing evidence and starts producing permission.
And there's one detail that sums it all up. A complaint like that one, AI will write, and write well. What it doesn't have is the person who made it. Nobody generates the one who confirmed an operation in the app, saw that it hadn't gone through, and never trusted that screen with their money again. That resentment was specific, it was theirs, and it was there that the information which changed the project was sitting.
Everything in its place. A synthetic user's place is not the user's chair.
A kiss, and see you next time!
Quick questions
What is a synthetic user in product research? It's asking a language model to answer as if it were your customer, generating a profile and interview answers without talking to anyone. It also shows up as synthetic persona or AI-generated participant.
Do synthetic users replace user research? As evidence for a decision, they don't. The available data shows minority, auxiliary use, and the authors of the study published in Nature Machine Intelligence urge caution whenever the identity of the participants matters to the task. As preparation for fieldwork, they help.
How do I decide whether to use a synthetic user in a given step? Ask who finds the error, and when. If the wrong answer surfaces next week, in a conversation with a real person, the risk is low. If it only surfaces at launch, don't use it.
What changed in the product after this research? The priority moved from features to trust, the redesign went across every platform, the app's rating rose from 3.5 to 4.7 on the App Store and Google Play, the target for new accounts was met, and clients who used to only check their balance started investing from their phones. Since the redesign touched several fronts at once, and other workstreams were running alongside, none of these numbers can be attributed to a single factor.
Why talk to a detractor? Because someone who is satisfied can't tell you what nearly went wrong. A detractor is someone scoring 0 to 6 on the Net Promoter Score recommendation scale, concentrates the churn risk, and describes the problem without diplomacy.
For anyone who wants the source:
- Bisbee, J., Clinton, J. D., Dorff, C., Kenkel, B. & Larson, J. M. (2024). Synthetic Replacements for Human Survey Data? The Perils of Large Language Models. Political Analysis, 32(4), 401–416 (read it on Cambridge). Open access. This is where the four findings come from: averages close to those of the 2016–2020 American National Election Study, less variation than the real survey, regression coefficients often differing, sensitivity to prompt wording, and different results from the same prompt across a three-month window. The model tested was GPT-3.5, with data collected in 2023.
- Wang, A., Morgenstern, J. & Dickerson, J. P. (2025). Large language models that replace human participants can harmfully misportray and flatten identity groups. Nature Machine Intelligence, 7, 400–411 (read it on Nature). The free version is the preprint accepted by the journal (arXiv:2402.01908), which carries the figures cited here: four models, 3,200 participants in total, sixteen demographic identities, and the inference-time techniques that reduce without removing the problem.
- User Interviews (2026). The State of Synthetic Users (read). Data collected from 11 to 22 May 2026, with 150 qualified responses plus five moderated interviews. A convenience sample recruited through the company's own channels, and the company sells human participant recruitment. A figure of 97% AI use circulates in third-party blogs and does not exist in the report: what's published is 80% regular use.
- Moran, K., Budiu, R., Gibbons, S. & The Experts at NN/g (2026). State of UX 2026: Design Deeper to Differentiate, Nielsen Norman Group, 16 January 2026 (read). The report doesn't cover synthetic users, and it's the source of the line about understanding users mattering more when there's AI in the product.
- Rosala, M. & Moran, K. (2024). Synthetic Users: If, When, and How to Use AI-Generated "Research", Nielsen Norman Group, 21 June 2024 (read). Where they test synthetic users against their own studies with real people and record the pull towards agreement, with the synthetic one appearing to care about everything.
- Reichheld, F. F. (2003). The One Number You Need to Grow, Harvard Business Review, December 2003. The origin of the Net Promoter Score, where the detractor is born. The 0-to-6 band and the link to churn are described by Bain & Company, which maintains the methodology (read).
- The client quotes here come from a research survey I ran in 2019, in internal project material, and were translated from Portuguese. The company's name stays out under a confidentiality agreement, and so do the internal base and conversion figures. The rating shift on both stores, from 3.5 to 4.7, and the new-accounts target are what I tracked during and after the redesign, and neither is a number that can be reconstructed from a public source today.
Read next

Too pretty to test: the aesthetic-usability effect
A beautiful interface makes users like your product more. It also makes them give you a high ease-of-use score on a test where they just struggled through the task right in front of you. The aesthetic-usability effect is the most polite trap in user research.
5 min read
Why you love one app and hate another (even when they do the same thing)
I'm paying off the promise from my last post: emotional design. Don Norman's three levels explain why we decide we like a product long before we understand it.
4 min read
Heads up: UX and UI design aren't the same thing — but they go hand in hand
UX and UI aren't the same role. Practical differences, where job posts blur the lines, and why both disciplines work together in product design.
3 min read
loading comments...