Two very different research tools are getting sold under the same name. One is AI-moderated interviewing, where an AI runs a dynamic conversation with real human participants. The other is synthetic users, where AI personas stand in for people entirely. They are not the same thing, they answer different questions, and the evidence for them is wildly different. Conflating them is how research teams end up making product decisions off data that looks real but is not.

The adoption is accelerating even though the evidence is uneven. Two in three researchers now use AI in at least some of their studies, up 19 points in a single year1. The demand side is moving just as fast: 66% report the volume of research demanded of them climbed, and the share of organizations where research is essential to all levels of strategy nearly tripled, from 8% to 22%1. More work, more authority, same headcount. That pressure is exactly why so many teams are grabbing at AI.

This article separates the two, shows what the 2026 evidence actually supports, and gives you a decision rule for when each one earns its place.

The distinction that matters

AI moderation is not a shortcut around people. It is a way to run more conversations with the same people, or with a wider pool of them, at a pace no human moderator can match. The standard definition, from platforms in this space, is "using artificial intelligence to autonomously conduct a dynamic conversation with a participant"2. The AI follows your interview guide, picks up on a vague answer, and decides to chase a thread or move on. The participant is a real, consenting human. You still need them, still recruit them, still owe them an incentive.

Synthetic users are the opposite bet. Here the participant is the model. You query an AI persona that was trained on behavioral data and ask it how a "user type" would react to a feature, a concept, or a set of tasks3. No recruiting, no no-shows, no panels. Answers in hours instead of weeks. The appeal writes itself.

The entire industry argument is about whether those two things are interchangeable. The 2026 evidence says they are not, and the gap is not subtle.

Two AI research paths: AI-moderated interviews run with real human participants and scale the signal you trust, while synthetic users substitute AI personas for people and only produce directional hypotheses
Two AI research paths: AI-moderated interviews run with real human participants and scale the signal you trust, while synthetic users substitute AI personas for people and only produce directional hypotheses

What AI genuinely does well in research

Start with where the expert guidance lands. Nielsen Norman Group reviewed which research tasks AI actually accelerates and reached a clear conclusion: it is most helpful in planning and analysis, and far less helpful in the moments that require watching a human do something4. AI can draft a study plan, generate and critique interview questions, produce recruitment emails and consent forms, take meeting notes, and later code transcripts and surface themes4.

What it cannot do is observe. In a usability test, an AI notetaker transcribes the words but misses what the user was doing, where they were on the page, which part of the screen they stared at, the hesitation before the click4. Behavioral data is about action, and a model that processes text does not see action. NN/g is blunt that tools which claim to conduct usability testing are overstating it: they can follow a script and respond to typed comments, but they cannot see where the participant is or what they are actually doing4.

That is the boundary line for the whole field. AI is excellent at the reading, summarizing, and pattern-finding parts of research. It is weak at the watching-parts. This holds for human participants, and it is the reason synthetic users fail at the same tasks from the other direction.

The synthetic users evidence is weak

When synthetic users started getting promoted hard, several vendors leaned on a single dramatic result. EY used one vendor's multi-agent simulation to replicate a 3,600-person global wealth survey and reported strong agreement across statistical metrics, completed in a day instead of six months5. That one anecdote has done enormous marketing work. The peer-reviewed record tells a more careful story.

MeasuringU reviewed 12 peer-reviewed papers comparing synthetic and human respondents across psychology experiments, surveys, and social research, and tallied the findings: 9 encouraging, 14 discouraging5. Read that again. The discouraging findings outnumber the encouraging ones.

The pattern is consistent, and it is the detail that is dangerous. Synthetic respondents match high-level means, then fail on the specifics. In one replication attempt, synthetic data matched average response levels but had inaccurate subgroup means, small standard deviations, and wrong regression coefficients5. In another, the expected personality factor structure came out but the item means deviated from real humans5. In the classic social-science replication test, only 3 of 14 studies reproduced, 5 failed, and 6 had data too uniform to even analyze5. Synthetic outputs cluster around the most common answer, which is precisely what makes them look clean and exactly what makes them useless for finding the unexpected.

The qualitative story is worse. In a CHI 2025 study, 19 UX researchers recreated a real human project using GPT-4. They were initially surprised by how plausible the narratives looked, then over several turns identified fundamental failures: responses lacked context and depth, and the exercise risked delegitimizing qualitative research5. A systematic literature review of 182 papers reaches the same structural conclusion: large language models are predictors of plausible text, not beings with lived experience, and that difference produces documented gaps between synthetic and human responses6. One researcher put the concern bluntly, calling synthetic data "inauthentic and similar to fake Amazon reviews, unrepresentative of real consumer opinions"6.

The people-pleasing problem deserves special attention. Synthetic users in one study claimed perfect online course completion rates. Real participants told messy stories about competing priorities, bad weeks, and shifting circumstances3. That messy part is where design insight usually lives. A participant who abandons a checkout because their child interrupted, or who cannot read the screen in bright sunlight, is telling you something a frictionless simulation will never produce. Synthetic users follow the most logical path to completion every time, and in doing so they miss exactly the surprises that reveal real flaws3.

Researchers know, even as they try it

Practitioners are not naive about this. In a May 2026 survey of 150 researchers, 47% said they were skeptical of synthetic users and wanted more evidence before trusting them, while 24% were cautiously optimistic6. Seventeen percent were outright opposed6. The top three concerns are all about validity: 88% worried about the quality and accuracy of insights, 79% worried stakeholders would overtrust AI-generated findings, and 79% worried about amplifying bias against underrepresented groups6.

The telling detail is that this caution coexists with heavy use. The same report found 80% of its respondents use AI regularly in their research work6. They have not rejected AI. They have rejected the idea that an AI persona is a substitute for a human participant. That is a sophisticated position, and it is the right one.

There is also a governance vacuum. More than 62% of researchers said their teams have no guidance at all on synthetic users, leaving the call to individual judgment6. That is a leadership opening for research teams, and a risk for everyone else, because a tool this seductive and this weakly validated should not be running ungoverned.

Where synthetic users actually earn their place

None of this means synthetic users are worthless. The honest read of the evidence, including from researchers who use them, is that they are good at the cheap, early, directional work and bad at anything that decides something real3.

The legitimate uses cluster at the front of a project. Before you commit to primary research, a synthetic persona can help you explore a problem space, surface assumptions worth testing, and generate questions you had not thought to ask. It can stress-test an interview guide or a survey flow for structural problems before you spend on real recruiting. It can scaffold a proto-persona for a workshop that you then validate against real people. For well-understood usability principles, clarity, navigation logic, cognitive load, it can give an early signal worth having3, the same discipline we walk through for keeping design systems authoritative through AI-generated code.

The disqualifying uses are the ones where the outcome has teeth. Concept testing is high on the list, because a synthetic persona validates whatever you put in front of it, and the critical friction that real participants add is exactly what you need to hear before committing to a direction3. High-stakes decisions, anything that touches protected groups or underserved populations, and any question where you do not already know the answer are all cases where the evidence says a synthetic persona is a hypothesis generator, not a source of truth. That matters more as the 2027 accessibility deadline forces more teams to test with assistive-technology users, exactly the kind of underserved population a synthetic persona cannot represent. One senior researcher made the risk concrete: if you rely on synthetic findings and they turn out not to represent the world, the only way to check is to launch the thing and see, which is a high-risk way to find out6.

The most defensible framing, from practitioners who both use and resist these tools, is that synthetic findings are hypotheses, not validated insights. Organizations that treat them as the latter, especially in low-maturity teams where the tool looks like a welcome shortcut, compound errors instead of catching them3.

Synthetic users decision rule: skip them for high-stakes or protected-group work, and use them only for pre-research hypothesis generation, guide testing, proto-personas, and screeners, then validate with real humans
Synthetic users decision rule: skip them for high-stakes or protected-group work, and use them only for pre-research hypothesis generation, guide testing, proto-personas, and screeners, then validate with real humans

A decision rule for your team

Here is the operating rule worth stealing. AI moderation earns its place in the reading and running parts of research: it conducts more interviews with real people, drafts and critiques your guides, transcribes, codes, and synthesizes. Use it to scale the human signal you already trust. Synthetic users earn their place only in the pre-research margin: hypothesis generation, guide piloting, screeners, proto-personas. Never use them to answer a question that will change a roadmap, reach an underserved group, or decide whether a concept ships. When in doubt, spend the recruiting budget.

The guardrails follow from that split. Before any synthetic use, run a risk assessment: is this a high-stakes decision, will the insights touch protected classes, do we know the answer already? If the use is exploratory, follow it up with human testing before anything depends on it6. Treat the output as directional, label it as synthetic in your repository, and never let a stakeholder mistake it for real research. A finding that cannot trace back to a human who actually used your product is a hypothesis with a nice suit on.

The deeper point is about what AI changed. It made the execution cheaper and faster, and that raised the bar on judgment rather than giving teams an excuse to skip research. The teams that will thrive in 2026 are the ones building repeatable processes for continuous learning and designing intentional AI-human workflows1. The AI handles the volume. A human still decides what it means, and whether the evidence is good enough to act on. That last job never got automated, and the research teams that pretend it did are the ones who will find out the hard way.

Sources

  1. Maze, The Future of User Research 2026. maze.co 2 3

  2. User Interviews, The Early Adopter's Guide to AI Moderation in UX Research. userinterviews.com

  3. User Vision, Synthetic Users and Digital Clones: A UX Researcher's Honest Take. uservision.co.uk 2 3 4 5 6 7

  4. Nielsen Norman Group, Accelerating Research with AI. nngroup.com 2 3 4

  5. MeasuringU, A Review of Experiments with Synthetic Users. measuringu.com 2 3 4 5 6

  6. User Interviews, The State of Synthetic Users (May 2026). userinterviews.com 2 3 4 5 6 7 8 9