5 Oct 2026, Mon

Can AI Stand In for Human Survey-Takers? A Pew Research Center Methodology Report

WASHINGTON — As artificial intelligence rapidly permeates every corner of modern industry, researchers are increasingly tempted to turn to large language models (LLMs) to cut costs, accelerate timelines, and bypass the logistical hurdles of polling human populations. But a sweeping new methodological study from the Pew Research Center suggests that while "silicon samples" and AI-generated "digital twins" are technologically fascinating, they are not yet ready to replace human respondents.

The comprehensive report, titled “Can AI Stand In for Human Survey-Takers? Not Really,” compares synthetic polling results directly against benchmark surveys administered to actual human respondents across three distinct waves of the Pew Research Center’s American Trends Panel (ATP). By building a custom architecture to administer synthetic surveys, researchers sought to determine whether AI models could successfully role-play as specific human beings—and whether synthetic data could reliably mirror human opinions on politics, social values, and global attitudes.

The findings offer a nuanced look at the state of AI in public opinion research. While advanced models like Anthropic’s Claude Opus 4.6 performed better than cheaper counterparts, even the most sophisticated configurations struggled to replicate the complexities of human sentiment, yielding average absolute survey errors that highlight significant boundaries in current LLM capabilities.


Main Facts: The Anatomy of a Synthetic Survey

To test the viability of AI polling, Pew Research Center researchers bypassed the traditional method of asking LLMs to generate aggregate point estimates for an entire population. Instead, they pioneered a "digital twin" approach.

In this framework, every human respondent from selected waves of the American Trends Panel was paired with a corresponding AI doppelganger. These synthetic respondents were fed a rich profile of demographic and psychological data, then prompted to complete the entire survey one question at a time—just as a human panelist would.

Key structural elements of the study included:

  • The Digital Twin Approach: Rather than producing generic, averaged macro-estimates, individual AI respondents answered questionnaires independently. This preserved internal consistency, allowed for flexible subgroup analysis (such as breaking down results by age, race, gender, or partisanship), and made it possible to apply standard survey weights.
  • The LLM Survey Pipeline: Synthetic respondents encountered the exact same programming as their human counterparts. If a panelist took a split-form survey, answered in Spanish, or faced a specific randomized question order, their digital twin experienced the exact same conditions. Furthermore, previous answers were fed back into the prompt to ensure internal consistency throughout the questionnaire.
  • Conditioning Information: The researchers tested varying levels of profile depth, ranging from basic demographic variables (age, education, race, gender, and political affiliation) to an "extended profile" incorporating 71 additional variables from the 2025 political typology survey.
  • Expert Reflection: To capture nuanced dispositions not explicitly stated in raw data, researchers utilized an automated "expert reflection" step. Before a survey wave was run, GPT-5.1 was prompted to adopt the persona of a political scientist and synthesize observations about each respondent’s underlying behavioral and political patterns.

Chronology: How the Experiment Was Conducted

The research was executed in a systematic, tournament-style testing framework. Because testing every possible combination of model, reasoning level, and profile depth was computationally and financially prohibitive, the Pew Research team evaluated dimensions sequentially, allowing the winning configurations to advance to subsequent rounds.

Phase 1: Initial Testing and Reasoning Level Evaluation

Initial testing began using GPT-5.1, OpenAI’s flagship model available at the time. Researchers first compared low, medium, and high model reasoning levels using a random subsample of 237 respondents from ATP Wave 185, paired with standard demographic profiles.

The results were unexpected: increasing reasoning effort drastically slowed down survey completion times without improving accuracy. Average person-level and question-level match rates hovered around 54% to 55% across all three reasoning levels, with virtually no meaningful variance among demographic subgroups. Finding no justification for the added time and financial expense, researchers selected low reasoning for all subsequent replications.

Phase 2: Evaluating Conditioning Information

Next, researchers tested the impact of richer profile data using the full sample of 8,512 respondents who completed Wave 185. They measured performance using the survey average absolute error—the average absolute percentage point error between the overall weighted estimate and actual ATP benchmark data.

  • Standard Profile Only: Yielded a survey average absolute error of 16 percentage points.
  • Extended Profile (adding 71 political typology variables): Improved performance significantly, shrinking the error to 13.7 percentage points.
  • Extended Profile + Expert Reflection: Further reduced error to 13.1 percentage points, confirming that deep behavioral context and expert political scientist summaries meaningfully improved model alignment.

Phase 3: Model Selection Tournament

With reasoning level and profile inputs standardized, researchers tested three distinct off-the-shelf closed models: OpenAI’s GPT-5.1 and GPT-5 nano, and Anthropic’s Claude Opus 4.6. Because Opus 4.6 carried a significantly higher financial cost, evaluations across all three models were conducted using a meticulously balanced subsample of 6,700 panelists from Wave 185.

The model comparison clearly demonstrated that intelligence architecture and size heavily dictate performance:

  • GPT-5 nano: This lightweight, budget-friendly model ran quickly but faltered in intelligence benchmarks. It recorded a survey average absolute error of 17.4 percentage points, performing worse than GPT-5.1 even when the latter was restricted to basic demographic profiles.
  • GPT-5.1: Serving as the industry baseline, this mid-priced model achieved a solid average absolute error of 13.3 percentage points on the subsample.
  • Claude Opus 4.6: Priced roughly five times higher than GPT-5.1, Anthropic’s state-of-the-art reasoning model scored highest on independent intelligence benchmarks and consistently outperformed its OpenAI competitors, registering a survey average absolute error of 11.4 percentage points.

Based on these outcomes, the final optimal configuration was established: Claude Opus 4.6 running at low reasoning, supplied with extended profile variables and expert reflections.


Supporting Data: Benchmark Results Across Survey Waves

Having established the optimal configuration, the research team deployed it across three distinct waves of the American Trends Panel to evaluate generalizability across different topics and sample sizes.

ATP Wave Topic Field / Replication Dates Respondents Replicated Number of Questions Model Reasoning Level Conditioning Information Average Absolute Survey Error (pct. pts.)
185 Politics Feb. 3 – Feb. 6 8,512 119 GPT-5.1 Low Standard profile 16.0
185 Politics Feb. 3 / Feb. 9 237 119 GPT-5.1 Medium Standard profile N/A
185 Politics Feb. 3 / Feb. 9 237 119 GPT-5.1 High Standard profile N/A
185 Politics Feb. 10 – 12 8,512 119 GPT-5.1 Low Extended profile 13.7
185 Politics March 2 – 3 8,512* 119 GPT-5.1 Low Extended profile + expert reflection 13.1
185 Politics March 12 – 13 8,512** 119 GPT-5 nano Low Extended profile + expert reflection 17.4
185 Politics March 9 – 12 3,200 119 Opus 4.6 Low Extended profile + expert reflection 11.4
185 (Cont.) Politics April 7 – 10 3,500 119 Opus 4.6 Low Extended profile + expert reflection —
190 Global Attitudes March 30 – April 2 3,398 105 Opus 4.6 Low Extended profile + expert reflection 14.7
192 Politics April 27 – May 1 4,981 67 Opus 4.6 Low Extended profile + expert reflection 11.2

(Notes: GPT-5.1 with extended profile/reflection was evaluated against other GPT-5.1 configurations using the full N=8,512 sample, but compared against GPT-5 nano and Opus 4.6 using the N=6,700 subsample. *GPT-5 nano was run on the full N=8,512 sample but evaluated against peers using the N=6,700 Opus 4.6 subsample subset.)

The data reveals that even under ideal conditions—leveraging state-of-the-art models like Claude Opus 4.6 backed by deep behavioral dossiers—synthetic surveys still produced average absolute errors ranging between 11.2 and 14.7 percentage points depending on the wave and subject matter. Performance was noticeably stronger on domestic political questionnaires (Waves 185 and 192) compared to complex global attitude assessments (Wave 190).


Official Responses and Methodological Insights

The findings strike a cautionary note for the data science and polling communities, pushing back against over-optimistic claims that synthetic populations can effortlessly substitute for human focus groups and electorate samples.

Researchers emphasized that while aggregate estimation approaches can sometimes mask underlying errors via the "averaging effect" of LLMs, building digital twins exposes the true granularity of where models succeed and fail. By examining individual responses, Pew researchers discovered that LLM-driven personas often struggled to authentically replicate the nuanced cognitive biases, inconsistencies, and emotional volatility inherent in real human decision-making.

Furthermore, the prompt engineering required delicate balance. The system prompts explicitly instructed the models not to answer with idealized, hyper-rational perfection, permitting synthetic respondents to exhibit incomplete knowledge, misinterpretations, or opt-outs due to survey fatigue or sensitive subject matter. Despite these safeguards, the residual error margins underline that an AI simulation remains fundamentally a statistical approximation rather than a living human mind.


Implications for the Future of Public Opinion Research

The implications of this comprehensive methodology report stretch across political science, market research, and commercial polling:

  1. Cost vs. Fidelity Tradeoffs: While running synthetic surveys is faster and less expensive than tracking down thousands of human panelists over weeks of field operations, the resulting data carries a baseline error rate that makes it risky for high-stakes forecasting, such as pre-election polling.
  2. The Importance of Rich Conditioning: The study conclusively proves that raw demographic data (age, race, gender) is entirely insufficient for creating realistic AI personas. To achieve even moderate alignment, synthetic respondents require extensive behavioral baselines, political typologies, and qualitative summaries (such as expert reflections).
  3. Model Selection is Critical: Organizations dabbling in synthetic research cannot rely on lightweight or cost-cutting models without severe penalties in accuracy. High-tier reasoning models like Anthropic’s Claude Opus 4.6 are mandatory to approach acceptable error thresholds.
  4. Complementary, Not Replacement: Ultimately, Pew Research Center’s findings suggest that silicon samples should not be viewed as standalone substitutes for human respondents. Instead, their highest and best use may lie in exploratory research, questionnaire pre-testing, counterfactual simulation, and understanding complex behavioral trends—leaving the measurement of actual public opinion firmly in the hands of living human populations.

By Basiran