Use CasesInsightsBlog
Alle Beiträge
Blog

Synthetic Personas: Plausible, but Not Representative

AI-generated personas give plausible answers and get the broad direction of survey results right. Three recent studies show what they miss: the variety, disagreement and contradictions of real people.

Von Johannes Tomin6 Min. Lesezeit
TeilenLinkedIn

The Promise... and the Problem

Synthetic respondents offer an almost irresistible proposition for market research. Instead of recruiting hundreds or thousands of people, you create AI-generated personas and ask them the same questions. There is no recruiting, no panel provider, no incentives and almost no fieldwork. Research that normally takes weeks could potentially be done in hours, at a fraction of the cost.

And at first glance, it works surprisingly well. Synthetic respondents give plausible answers and can reproduce many of the broad patterns we see in real research.

But plausible is not the same as representative. Real people are messy. People who look very similar on paper can think and behave very differently. They contradict themselves, have unusual combinations of attitudes and don't always fit neatly into the persona we would expect them to be.

LLMs tend to smooth out exactly this messiness. They are good at creating believable people, but those people are often more similar, more consistent and more predictable than real ones.

The central question is therefore not whether an LLM can give a believable answer on behalf of a fictional customer. It clearly can. The question is whether a thousand fictional customers, taken together, answer the way a thousand real customers would: with the same range of opinions, the same disagreements and the same surprises.

The evidence so far suggests they don't. Synthetic personas get the broad direction right: they know which groups tend to be more satisfied, more engaged or better off. But they disagree far less than real people do, their opinions fall into neater patterns, and the surprises are mostly missing.

Synthetic Respondents

At the centre of most synthetic respondent approaches is the persona: a description of a fictional person, based on characteristics such as age, background, attitudes and behaviour.

A simple persona might contain only age, gender and location. A richer one can add income, education, occupation, attitudes, personality, political views, hobbies, habits or even an entire life story. The AI is then asked to answer survey questions as that person.

On the surface, this sounds straightforward. Create enough different personas, ask each of them the same questions, and you have something that looks like a synthetic population. The individual answers can be surprisingly plausible. But that raises a much more important question: does giving an AI a persona actually make it behave like a real person from that population?

That is a very different test from asking whether an AI can produce a believable answer. A synthetic population needs to reproduce not just plausible individual responses, but also the variation, disagreement, inconsistencies and relationships between answers that exist in real populations.

Where the Synthetic Population Falls Apart

Recent research suggests that this is where the problems begin. In one study, researchers tested 37 LLMs and generated around 65,000 synthetic questionnaires, comparing the results with responses from real people. The findings reveal a recurring pattern: the models can reproduce some broad characteristics of human responses, but they systematically distort how people differ from one another.

1. The models agree with each other more than with real people.
Across 37 different AI models, the models' answers resembled one another more than even the best model resembled real respondents. Instead of 37 genuinely different simulated populations, you effectively get 37 variations of a similar idea of how people behave. Averaging all 37 models made the results worse, not better.

2. Understanding language is not the same as understanding populations.
The best of the 37 models was only slightly better at reproducing the overall pattern of real answers than a simple statistical method with no AI at all. Both fell well short of how closely two groups of real people resembled each other. In other words, being good at understanding questions and generating plausible language does not automatically make a model good at reproducing population-level behaviour.

3. The answers are too consistent.
A persona gives the model a coherent story about who someone is. The model then tends to make the person's answers fit that story. Related answers become unusually consistent with one another, making the synthetic respondent look like a remarkably coherent human being. Real people are messier. Their opinions do not always line up neatly, and their answers to related questions can contain contradictions. Synthetic respondents tend to smooth out this noise, which can make relationships between variables look stronger and cleaner than they really are.

4. The answers are too agreeable.
Synthetic respondents also tended to agree more often than real people, avoid strong positions and gravitate towards the middle of the answer scale. The result is a synthetic population that is more polite, more consistent and more agreeable than the real one. These effects matter because they are not just cosmetic differences. Once synthetic responses are aggregated, the distortions can start looking like real statistical relationships.

5. The AI invents relationships.
The researchers tested ten chains of the form “X affects Y through Z” where no such effect existed in the real data. In three cases, the synthetic data nevertheless showed a statistically significant effect. The authors conclude that some relationships reproduced by the AI are “partly confabulated, not purely recovered” and warn against using synthetic respondents to discover new cause-and-effect relationships.

6. Demographics can trigger stereotypes.
Changing only one characteristic of a persona could substantially change its answers. The strongest effect came from education, with much smaller effects for job role and gender. Some of these effects, including gender differences, did not exist in the real data at all. The persona therefore does not simply describe a respondent. Its characteristics can actively shape what the model expects that respondent to say.

7. More synthetic respondents do not fix the problem.
If 100 synthetic respondents are not representative, generating 100,000 does not make them representative. It only gives you a much more precise estimate of the same distorted population. More data reduces random sampling error. It does not remove systematic bias in the process that generated the data.

At this point, it would be tempting to conclude that synthetic respondents are simply useless. The evidence does not support that conclusion either.

8. The AI is still useful, within limits.
The models often reproduce the direction of relationships found in real data: if two things go together in reality, the synthetic data usually shows the same pattern. They can also distinguish some groups correctly and recognise which survey questions belong together. What they often get wrong is the strength of those relationships and the underlying structure.

That distinction is crucial. An AI can be very good at producing a plausible answer from a persona without producing a statistically faithful representation of the population that persona is supposed to represent.

So the real question is not simply “Can synthetic respondents replace real ones?” It is “Where can we use AI-generated respondents without confusing plausible answers with evidence about real people?”

If you want to dig deeper into synthetic personas, the evidence behind these findings, and what it means for using AI as a substitute for real respondents, I’ve explored the topic in more detail on Substack.

Sie wollen wissen, was Ihre Kunden wirklich denken?

Sprechen Sie uns an