6 Oct 2026, Tue

Can Artificial Intelligence Stand In for Human Survey-Takers? Pew Research Center Experiment Says “Not Really”

As artificial intelligence rapidly infiltrates various sectors of data science and market research, a provocative question has emerged: Can AI-generated synthetic data accurately mirror the results of high-quality public opinion polls?

To find out, the Pew Research Center conducted an extensive methodological experiment. Simulating thousands of real human respondents using advanced large language models (LLMs), researchers discovered that AI falls short of reliably duplicating human public opinion. The choice of the AI model alone creates starkly contradictory portraits of the American public, proving that silicon samples cannot yet replace direct engagement with the public.


Main Facts and Overview

The Pew Research Center’s comprehensive study evaluated whether "digital twins"—AI models programmed to adopt the personas of real human panelists—could replicate actual public opinion surveys. Utilizing detailed demographic profiles and historical survey responses, researchers tested OpenAI’s GPT-5.1 and Anthropic’s Claude Opus 4.6 against identical questionnaires previously completed by human members of the American Trends Panel (ATP).

AI polling results change based on the model used

The core takeaway is definitive: neither model accurately reflected actual public opinion.

  • High Error Margins: When examining survey average absolute error, Claude Opus 4.6 exhibited an average absolute error of 11.4 percentage points, while OpenAI’s GPT-5.1 fared worse at 13.3 percentage points.
  • Contradictory Narratives: Each synthetic sample painted a radically different picture of the American public. GPT consistently exaggerated public dissatisfaction and leaned toward extreme response options (such as "always" or "never"), while Opus tended to flatten ideological intensity, pulling responses toward moderate, middle-of-the-road choices.
  • The Model Choice Crisis: Depending entirely on which AI model was selected, analysts would draw fundamentally opposing conclusions about crucial topics like national direction, abortion legalization, and partisan political loyalty.

Chronology of the Experiment

The methodological experiment was structured and executed across several distinct phases during late 2025 and the first half of 2026:

  1. Late 2025 (Foundation Building): Researchers gathered baseline data by administering comprehensive questionnaires, including a 2025 political typology survey, to the human participants of the American Trends Panel (ATP).
  2. January 20–26, 2026 (Human Baseline Survey): The Center conducted a primary public opinion survey of U.S. adults to capture fresh human sentiment on governance, economic concerns, and political values (Wave 185).
  3. March 2–3, 2026 (First AI Replication): Using a "digital twins" approach, OpenAI’s GPT-5.1 was provided with extended profile information for 6,700 ATP panelists. The model was instructed to answer Wave 185 questionnaires sequentially, following the exact programming of the human survey.
  4. March 9–12 and April 7–10, 2026 (Second AI Replication): Anthropic’s Claude Opus 4.6 underwent the exact same simulation protocol to generate a parallel synthetic dataset for Wave 185.
  5. Spring 2026 (Additional Waves): Additional synthetic tests were run across subsequent human survey waves, specifically Wave 190 (focusing on the U.S. role in the world) and Wave 192 (addressing national problems).
  6. September 30, 2026 (Public Release): Pew Research Center published its comprehensive findings, data labs breakdown, and methodological assessments under the title "Can AI Stand In for Human Survey-Takers? Not Really."

Supporting Data and Comparative Analysis

To understand the depth of the discrepancies, Pew analysts broke down the synthetic results across multiple sensitive socio-political domains, revealing deep systemic biases within the AI models.

AI polling results change based on the model used

1. Public Views of Politics and Democracy

Long-standing trend questions regarding the country’s direction exposed wild inconsistencies between the models:

  • Dissatisfaction with the Country’s Direction: While human respondents reported a 69% dissatisfaction rate (closely matched by Opus at 70%), GPT estimated that a staggering 100% of Americans were dissatisfied.
  • Political "Sides" Winning or Losing: Human polls indicated that 63% felt their side of politics had been losing more often than winning. Opus matched this at 63%, but GPT once again inflated the sentiment to 100%.
  • Solutions to Big Issues: When asked whether there are clear solutions to major national issues, 56% of humans agreed. GPT closely mirrored this at 57%, but Opus severely undercounted this sentiment, dropping it down to 25%.

2. Abortion Attitudes

Abortion policy highlights how neither model captures the full scope of public sentiment:

  • Human Baseline: Roughly 60% of Americans stated abortion should be legal in all or most cases, while 39% said it should be illegal.
  • Opus Synthetic Results: Opus produced a similar aggregate split (62% legal, 37% illegal) but severely underestimated the share of Americans holding staunch, absolute views.
  • GPT Synthetic Results: GPT estimated an even split that leaned slightly toward illegality (48% legal vs. 53% illegal), failing to capture the broader pro-legalization majority seen in human data.

3. Partisan Attitudes Among Republicans

Evaluating the political right demonstrated how AI model choice distorts intra-party dynamics:

AI polling results change based on the model used
  • Trump Approval: Among 2024 Trump voters, 65% approved very strongly, and 19% approved not so strongly in human responses. Opus reversed this dynamic, estimating that 53% approved not so strongly and only 46% very strongly. GPT swung in the opposite direction, pushing "very strong" approval to 88%.
  • China’s Relationship: Asked to characterize China, human respondents split between "Competitor" (60%) and "Enemy" (28%). Opus leaned heavily into "Competitor" (85%), while GPT inflated the "Enemy" category to 44%.
  • Congressional Loyalty: Asked if Republican lawmakers had an obligation to support Donald Trump’s policies even if they disagreed, 61% of human Republicans said they do not have that obligation. Opus radically overestimated independence at 81%, whereas GPT concluded that 73% did feel an obligation of absolute loyalty.

4. Extreme vs. Moderate Response Bias

A pervasive trend across the experiment involved response scaling:

  • GPT’s Extreme Bias: GPT systematically gravitated toward the extreme ends of rating scales (such as "always," "never," "extremely," or "not at all"), artificially amplifying perceived societal polarization.
  • Opus’s Moderate Bias: Conversely, Opus favored safe, middle-of-the-road answer options (such as "somewhat," "about right," or "neither"), smoothing out ideological edges and erasing strong convictions.

Official Stance and Policy of Pew Research Center

The motivation behind this sweeping experiment was academic rigor, not a pivot in corporate strategy. As synthetic polling grows increasingly popular across the commercial insights industry, researchers felt an urgent need to evaluate the underlying validity of silicon samples.

However, the Center has drawn a hard line regarding its own operations: speaking directly to the public remains essential to measuring true public opinion.

AI polling results change based on the model used

Pew Research Center has explicitly stated that it has no current or future plans to use AI models to generate survey results. This methodological experiment was strictly designed to audit and stress-test an emerging industry trend rather than validate its adoption.


Implications for the Future of Polling

The implications of Pew’s experiment send a sobering warning to market researchers, political strategists, and news organizations tempted by the low cost and high speed of AI-generated survey data.

  1. The Illusion of Reliability: Because LLMs are trained on vast corpuses of human text, they can easily mimic the tone and syntax of survey responses. However, mimicking human voice is not the same as mirroring human psychology or shifting sentiment.
  2. Vulnerability to Model Drift and Bias: As demonstrated by the divergence between GPT-5.1 and Claude Opus 4.6, synthetic data is entirely downstream of the specific corporate architecture, training data, and alignment guardrails of the underlying AI model. Choosing a different vendor fundamentally changes the socio-political reality reported by the data.
  3. Erosion of Public Trust: If campaigns, media outlets, and policymakers begin relying on synthetic samples that artificially inflate anger (like GPT) or artificially suppress passion (like Opus), public policy and media narratives risk being built on hallucinations rather than reality.

Ultimately, Pew’s research delivers a definitive verdict: while artificial intelligence will continue to transform data processing and qualitative analysis, it cannot currently stand in for the complex, unpredictable, and deeply human experience of answering a public opinion poll.