What the research
actually supports.
Simulated respondents are a real research method with a real literature — including the studies that show where it fails. This page collects both, and states plainly what we do and don't claim.
We don't publish a parity score
Several vendors advertise a single accuracy figure against real research — “90% match”. We don't, because we haven't run a benchmark that would justify one, and a number produced on someone else's study population wouldn't transfer to yours anyway. A parity claim is only meaningful alongside the concept, the audience and the questions it was measured on.
What we do instead is show the working. Every figure in a report is labelled measured — counted in code from stored participant answers — or interpreted, written by the analyst model. Every study carries a credibility band, a margin at 95% confidence, a saturation curve, quote-grounding checks and a full audit trail. You can reproduce the measured layer yourself from the same stored responses.
If you run a synthetic study and a real one on the same concept, we'd genuinely like the comparison. That is how a parity claim gets earned.
The value is in the deviation
A synthetic panel that agreed perfectly with real fieldwork would tell you nothing you couldn't get by waiting six weeks. The useful signal is where a simulated audience reacts differently from what you expected — an objection you hadn't priced in, a segment that splits, a price point where enthusiasm falls off a cliff.
Treat the output as a hypothesis generator with a ranking attached. It tells you which questions are worth asking real people, and which options aren't worth testing at all.
Ranking concepts, price points or messages against each other.
Surfacing objections and framing before you write the discussion guide.
Forecasting demand, market share or absolute conversion rates.
What we borrowed from the literature
Individual conditioning, not one summary
Each participant is generated with a full backstory and interviewed separately — the algorithmic-fidelity finding is about conditioning on a specific person, not asking a model for an average.
Stable personality per participant
A fixed Big Five profile per persona, carried into every answer, keeps individuals distinct instead of collapsing into a single agreeable voice.
Saturation rather than sample-size theatre
We track when additional participants stop introducing new themes, which is the qualitative-research answer to “is this enough?”.
Dispersion reported, not hidden
Because simulated answers under-disperse, we report spread and consensus alongside the mean, and show ranges where the cell is thin.
Named limits on thin audiences
Model opinion coverage is uneven across demographics. Vague or niche audience specs are the weakest case, and the credibility checks say so.
Where the method holds up
- Out of One, Many: Using Language Models to Simulate Human Samples
Argyle, Busby, Fulda, Gubler, Rytting, Wingate · Political Analysis, 2023
Introduces “algorithmic fidelity”: conditioning a model on a detailed backstory reproduces response patterns of the corresponding human subgroup.
- Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies
Aher, Arriaga, Kalai · ICML, 2023
“Turing Experiments” replicate classic behavioural studies with simulated participants, and document where the replication breaks.
- Generative Agents: Interactive Simulacra of Human Behavior
Park, O'Brien, Cai, Morris, Liang, Bernstein · UIST, 2023
Shows that agents with stable memory and profile produce believable, individually distinct behaviour rather than one averaged voice.
- Generative Agent Simulations of 1,000 People
Park, Zou, Shaw, Hill, Cai, Morris, Willer, Liang, Bernstein · arXiv, 2024
Interview-grounded agents of 1,052 real participants; reports how closely agents match their source person's survey answers.
- Large Language Models as Simulated Economic Agents (homo silicus)
Horton · NBER Working Paper, 2023
Argues simulated agents are cheap pilot subjects: run the experiment in silico first, then spend money on the version worth running.
Where it breaks
These are the papers a sceptical colleague will bring to the meeting. We'd rather hand them over ourselves.
- Whose Opinions Do Language Models Reflect?
Santurkar, Durmus, Ladhak, Lee, Liang, Hashimoto · ICML, 2023
Model opinion distributions skew toward some demographics and away from others. Under-represented groups are the least reliable to simulate.
- Synthetic Replacements for Human Survey Data? The Perils of Large Language Models
Bisbee, Clinton, Dorff, Kenkel, Larson · Political Analysis, 2024
Simulated responses show lower variance and unstable correlations versus real survey data — direction can hold while dispersion does not.
- Can Large Language Models Transform Computational Social Science?
Ziems, Held, Shaikh, Chen, Zhang, Yang · Computational Linguistics, 2024
Useful as zero-shot annotators and hypothesis generators; not a substitute for the human data they are benchmarked against.
