Pew finds AI models are poor stand-ins for human survey respondents
Testing Claude Opus 4.6 and GPT-5.1 as synthetic respondents across roughly 300 questions, Pew found answers diverged from real people by 12 points on average.
- Culture & impact
- Benchmarks & progress
- Minor
Pew Research Center published a study testing whether large language models can stand in for real survey respondents, and concluded they generally cannot. Researchers built “digital twins” — synthetic respondents profiled to resemble members of Pew’s American Trends Panel — using Anthropic’s Claude Opus 4.6, which Pew said was the best-performing model of several it evaluated, and compared their answers against nearly 300 questions from surveys fielded across three waves between January and April 2026.
On average, the AI-generated estimates differed from the real human results by 12 percentage points, and errors exceeded 15 points on roughly 28% of questions. Mistakes were not evenly distributed: Black and Republican-leaning respondents were simulated less accurately than other groups, and the model badly overestimated public knowledge — 98% of synthetic respondents answered a basic First Amendment question correctly, against 52% of real respondents. On a subset of questions, Pew also tested OpenAI’s GPT-5.1 and found its errors ran the opposite way from Opus’s: the GPT estimates described a public with more extreme opinions than it actually held, while the Opus estimates described one more middle-of-the-road than reality.
Pew said the results “differ in ways that are often unpredictable” depending on the model and method used, and stated it had no plans to use AI to replace real survey respondents. The finding bore on a broader ambition — simulating public opinion cheaply at scale using language models — that the study found no model had yet demonstrated it could do reliably.
In the commentary
What people were saying around this time — external links, from the record's commentary rail.
- 29 September 2026 · Jasper Jackson · TransformerScrapping Astra 6.1 looks like a good call. OpenAI shouldn’t be the one to make it.
- 28 September 2026 · Zvi Mowshowitz · Don't Worry About the VaseWhat Also Happened: #NotOnlyHuggingFace
- 27 September 2026 · Zvi Mowshowitz · Don't Worry About the VaseThe Quest for Embedded Evaluators