The science

What the research
actually supports.

Simulated respondents are a real research method with a real literature — including the studies that show where it fails. This page collects both, and states plainly what we do and don't claim.

Our position

We don't publish a parity score

Several vendors advertise a single accuracy figure against real research — “90% match”. We don't, because we haven't run a benchmark that would justify one, and a number produced on someone else's study population wouldn't transfer to yours anyway. A parity claim is only meaningful alongside the concept, the audience and the questions it was measured on.

What we do instead is show the working. Every figure in a report is labelled measured — counted in code from stored participant answers — or interpreted, written by the analyst model. Every study carries a credibility band, a margin at 95% confidence, a saturation curve, quote-grounding checks and a full audit trail. You can reproduce the measured layer yourself from the same stored responses.

If you run a synthetic study and a real one on the same concept, we'd genuinely like the comparison. That is how a parity claim gets earned.

How each figure is produced
The interesting part

The value is in the deviation

A synthetic panel that agreed perfectly with real fieldwork would tell you nothing you couldn't get by waiting six weeks. The useful signal is where a simulated audience reacts differently from what you expected — an objection you hadn't priced in, a segment that splits, a price point where enthusiasm falls off a cliff.

Treat the output as a hypothesis generator with a ranking attached. It tells you which questions are worth asking real people, and which options aren't worth testing at all.

Strong use

Ranking concepts, price points or messages against each other.

Strong use

Surfacing objections and framing before you write the discussion guide.

Weak use

Forecasting demand, market share or absolute conversion rates.

Method

What we borrowed from the literature

  • Individual conditioning, not one summary

    Each participant is generated with a full backstory and interviewed separately — the algorithmic-fidelity finding is about conditioning on a specific person, not asking a model for an average.

  • Stable personality per participant

    A fixed Big Five profile per persona, carried into every answer, keeps individuals distinct instead of collapsing into a single agreeable voice.

  • Saturation rather than sample-size theatre

    We track when additional participants stop introducing new themes, which is the qualitative-research answer to “is this enough?”.

  • Dispersion reported, not hidden

    Because simulated answers under-disperse, we report spread and consensus alongside the mean, and show ranges where the cell is thin.

  • Named limits on thin audiences

    Model opinion coverage is uneven across demographics. Vague or niche audience specs are the weakest case, and the credibility checks say so.

Evidence base

Where the method holds up

Counter-evidence

Where it breaks

These are the papers a sceptical colleague will bring to the meeting. We'd rather hand them over ourselves.