How accurate are synthetic respondents? What the evidence says in 2026

Vendors quote 85 to 95 percent parity. Independent reviews say accuracy has no universal number. Here is how to read the claims, what calibration actually means, and the one test that settles it for your business.

Published 15 September 2026 · 2 min read · By the Precheck team

Key takeaways

  • Published parity figures range from roughly 0.7 to 0.9 correlation with human responses, and vary by question type, audience, and model.

  • Parity on a benchmark is not accuracy on your decision. The only number that matters is how close the model got on decisions like yours.

  • Calibration means running the model blind against a known outcome and measuring the gap. Ask every vendor for it.

  • Precheck has not established predictive accuracy. We are running a calibration round to measure it rather than claim it.

Every synthetic research vendor has an accuracy number. Synthetic Users cites 85 to 92 percent synthetic-organic parity. Aaru cites a 0.90 median correlation in an EY engagement. Academic work on simulated survey responses reports model-question correlations from about 0.70 to 0.92 depending on the pairing. Independent reviewers land on a less quotable conclusion: there is no universal accuracy rate.

Both things are true. The numbers are real measurements. They just do not measure what a buyer needs to know.

Parity is not accuracy

A parity score compares a synthetic response set with a human response set on the same instrument. It answers: does the model produce answers that look like the answers people give? It does not answer: did the people then do what they said? Aaru's own positioning leans on this gap. Its argument is that humans are unreliable narrators of their own behaviour, so matching stated preferences is the wrong target.

For a pricing or bundle decision, the target is observed behaviour: what share actually bought at ₹1,299, whether the bundle cannibalised singles, whether churn moved after the packaging change. A model can score 0.9 against a survey and still miss those.

What moves accuracy up or down

From the published work and our own experience, four things dominate.

FactorPushes accuracy upPushes it down
Underlying dataThe company's own customer data for that audienceGeneral-population training data only
Question typeRelative comparisons between scenariosAbsolute point estimates
NoveltyDecisions with close analogues in the dataGenuinely new products or categories
ValidationBlind runs against known outcomes in your categoryVendor benchmarks on public datasets

The pattern is consistent with what NielsenIQ and Greenbook describe: accuracy is earned per audience and per task, and it decays as the question moves away from the data.

The test that settles it

There is one procedure that turns a claim into a measurement, and it is not complicated.

  1. Pick a decision you already made: a launch, a price change, a bundle.
  2. Give the model the inputs you had at the time, and nothing you learned afterwards.
  3. Let it simulate the response.
  4. Compare the simulated distribution with what actually happened.
  5. Repeat across a handful of decisions and audiences.

The result is a number that belongs to you and your category. If a vendor will not run this, or will only show it on their own examples, that tells you something about the number they quote.

Precheck's position

We have not established predictive accuracy, and our site says so on the page rather than in a footnote. Simulations are underway for a small number of large companies on data they have shared. We are inviting teams to bring a decision they have already made, with the outcome, so we can run it blind and share the comparison. If the simulation is wrong, we learn where. If it is close, that is the first honest accuracy claim we will make.

Questions people ask

What does calibration mean in synthetic research?
Calibration is running a simulation blind against a decision whose real outcome is already known, then measuring how far the simulated response distribution sat from what happened. It is the difference between a claim and a measurement.
Are synthetic respondents accurate enough to replace surveys?
For exploration and prioritisation, often yes. For decisions where the exact number matters, such as a price you will print on a shelf, no evidence supports skipping the real-world check.
Why do accuracy figures vary so much between vendors?
Because they are measured on different questions, audiences, and data. A model that tracks stated preferences well can miss actual purchase behaviour badly. Compare vendors on your decision, not their benchmark.

Everyone

What is synthetic customer research? A plain-English guide

Synthetic customer research uses AI models to simulate how customers might respond to a product, price, or message. Here is what it is, how it works, where it breaks, and how to use it without fooling yourself.

15 September 2026 · 3 min read

Everyone

How to test a price change before you launch it

A practical sequence for D2C brands and SaaS teams: frame the decision, explore how each segment might respond, then run the one real-world test that settles it. Includes what to measure and the mistakes that make price tests worthless.

15 September 2026 · 3 min read