![]()
WASHINGTON — As artificial intelligence systems grow exponentially more sophisticated, the polling and survey research industry faces a profound technological crossroad. The central proposition is tempting for researchers and cost-conscious organizations alike: Instead of mobilizing vast resources to contact, interview, and survey thousands of everyday citizens, pollsters could theoretically prompt an advanced AI model to simulate how a given population would respond.
These simulated cohorts—often referred to in computational science as "synthetic samples" or "digital twins"—promise rapid results at a fraction of the time and financial expense of traditional polling. However, a major empirical study released by the Pew Research Center delivers a stark, definitive verdict on this emerging trend. According to the comprehensive methodological experiment, AI-generated survey data consistently fails to accurately duplicate the results of high-quality public opinion polls, exhibiting wild inaccuracies, a bias toward intellectual elitism, and a troubling tendency to lean into demographic stereotypes.
The Center’s leadership has made it clear that while AI has valid administrative functions in modern data science, speaking directly to the public remains an irreplaceable cornerstone of measuring genuine human sentiment.
Chronology of the Experiment: Testing the Digital Twin
To understand the capabilities and limitations of AI-based polling, Pew Research Center researchers designed a rigorous methodological framework during the first half of 2026. The goal was to test whether state-of-the-art language models could accurately act as human proxies when fed rich, detailed profiles of actual individuals.
The experimental timeline unfolded across several distinct phases:

- Laying the Foundation (2025–January 2026): Researchers utilized the American Trends Panel (ATP), a nationally representative, probability-based panel of U.S. adults. Panelists completed detailed demographic questionnaires and participated in the Center’s comprehensive 2025 political typology survey, providing deep baseline data on their political leanings, core values, and socioeconomic backgrounds.
- Wave 185 Administration (January 20–26, 2026): Human panelists completed the first target survey wave examining public attitudes and political perspectives. This wave was subsequently replicated using synthetic respondents between March 9–12 and April 7–10, 2026.
- Wave 190 Administration (March 23–29, 2026): Human participants tackled civic knowledge and institutional trust metrics. The AI replication of this wave was conducted immediately afterward between March 30 and April 2, 2026.
- Wave 192 Administration (April 20–26, 2026): The final target wave addressed pressing national problems and economic stressors. Synthetic execution of this wave occurred between April 27 and May 1, 2026.
- Model Selection and Execution: Although researchers tested multiple large language models, the primary synthetic results were generated using Anthropic’s Claude Opus 4.6, operating with extended profile information and expert reflection settings.
The results of this months-long cross-examination revealed profound discrepancies between what real humans think and what synthetic models believe they think.
Supporting Data: The Anatomy of Synthetic Error
Across nearly 300 individual survey questions spanning diverse topics, formats, and demographics, the estimates produced by the AI respondents deviated drastically from their human counterparts.
1. High Margins of Absolute Error
When comparing aggregate responses, the AI-generated estimates differed from human poll results by an average of 12.4 percentage points across all tested waves. Republicans and Republican-leaning respondents saw an average error rate of 16.1 percentage points, while Democrats and Democratic-leaning respondents experienced a 13.6-point average error. Racial and ethnic breakdowns yielded similarly high deviations: White respondents showed a 12.5-point average error, Hispanic respondents 13.9 points, Asian respondents 13.5 points, and Black respondents a striking 15.1-point average error.
On roughly 28% of all questions asked, the discrepancy between human and synthetic answers exceeded 15 percentage points. In several instances, individual answers missed the mark by 20, 30, or even 40 points.
2. Misses on Timely and Topical Issues
The AI models struggled immensely when tasked with capturing nuanced shifts in public opinion regarding unfolding current events. For example, while 38% of surveyed U.S. adults in early 2026 stated it was acceptable for immigration officers to wear face coverings, the synthetic model predicted that only 17% would agree—a massive 21-percentage-point underestimation.

Similarly, when asked if they had heard a lot about the rapid expansion of data centers, 25% of human adults answered in the affirmative, whereas the AI model predicted a nearly nonexistent 3% (a 22-point error). Conversely, the model overestimated presidential approval ratings for Donald Trump, recording a 46% synthetic approval rating compared to the actual human baseline of 34%.
3. Eradication of Extremes and Nuance
Human opinion is famously broad and multifaceted, yet AI models demonstrated a severe tendency toward homogenization. Nearly half of all tested questions contained at least one response option that was completely ignored—not selected by a single synthetic respondent.
This statistical flattening severely distorted contentious debates. On the issue of abortion access, for instance, human respondents displayed realistic polarization: 23% stated abortion should be legal in all cases, while 11% said it should be illegal in all cases. The AI model successfully captured the overarching majority sentiment regarding overall legality, but completely failed to capture the fringes, pegging the "illegal in all cases" response at 0% while overestimating moderate positions.
4. The Intelligence Paradox: AI Overestimates Human Knowledge
One of the most profound behavioral flaws uncovered in the study was the AI’s inability to accurately replicate human ignorance or uncertainty.
When tested on basic civic and international knowledge questions, the performance gap was staggering:

- First Amendment Guarantees: While 52% of U.S. adults correctly identified rights protected by the First Amendment, 98% of synthetic respondents answered correctly.
- NATO’s Central Focus: 56% of real humans answered correctly, compared to 99% of the AI model’s digital personas.
- Greenland’s Territory Status: 63% of U.S. adults knew Greenland’s geopolitical status, versus 82% of synthetic respondents.
Furthermore, when given the option to answer "not sure" or express genuine uncertainty, human panelists were roughly four times as likely to select it as the AI models. The silicon respondents preferred to project confidence rather than admit a lack of knowledge.
5. Model Dependency: Different Engines, Different Biases
The choice of underlying AI architecture radically alters the output. A comparative analysis between Anthropic’s Claude Opus 4.6 and OpenAI’s GPT-5.1 revealed that different models skew public opinion in entirely opposite directions. GPT-5.1’s synthetic estimates depicted an American public with significantly more extreme, polarized views than reality, whereas Claude Opus 4.6 painted a picture of a population that was far more moderate and centrist than actual survey data indicated.
Official Responses and Institutional Stance
The leadership at Pew Research Center has been unambiguous regarding the implications of these findings for the broader polling ecosystem.
"The Center believes that speaking to the public is essential to measuring public opinion and has no current or future plans to use AI models to generate survey results," the organization stated in its official methodology guidelines.
Methodologists emphasize that while artificial intelligence is a revolutionary tool for streamlining computational tasks, text analysis, and code generation, treating it as a cognitive mirror for humanity is fundamentally flawed.

"It’s not just that the AI results differ from human results, although that is certainly true," researchers noted in the study’s conclusions. "The bigger story is that these results differ in ways that are often unpredictable, and they are highly subject to factors like unforeseen real-world events or the choice of model used."
Broader Implications for the Future of Survey Research
The publication of "Can AI Stand In for Human Survey-Takers? Not Really" serves as a crucial reality check for an industry increasingly seduced by automation. As political campaigns, market researchers, and academic institutions look for ways to cut costs, the temptation to deploy synthetic samples will undoubtedly persist.
However, the downstream consequences of relying on AI-generated public opinion could be catastrophic for democratic integrity. If synthetic polls systematically underestimate extreme political viewpoints, erase marginalized perspectives, and falsely project universal civic literacy, decision-makers acting on those polls will operate in a fabricated reality.
AI models are trained on vast corpuses of human text, making them exceptional at mimicking the language of humanity, but utterly incapable of capturing the lived, emotional, and unpredictable nature of human thought. Until artificial intelligence can genuinely experience the anxieties of cost-of-living crises, the complexities of moral debates, and the genuine limits of civic education, survey-taking will remain firmly—and necessarily—the domain of real human beings.
