
WASHINGTON — In an era where artificial intelligence is rapidly reshaping industries from software engineering to creative arts, a new methodological experiment from the Pew Research Center explores whether advanced AI models can successfully replicate the gold standard of public opinion polling: human responses.
The results suggest that while AI can mimic the general contours of public sentiment, it falls short of capturing the rich, nuanced, and diverse tapestry of real human thought. According to the Pew report, AI-generated "synthetic samples" tend to flatten public opinion, stripping away the complex extremes and outliers that characterize real-world populations. Furthermore, AI models frequently avoid certain answer choices entirely, creating a distorted, one-note reflection of society.
Main Facts
The Pew Research Center’s landmark methodological experiment investigated whether synthetic survey data generated by an AI model could accurately duplicate the results of high-quality public opinion polls.
To conduct the study, researchers utilized a "digital twins" approach. They tasked Anthropic’s Claude Opus 4.6—at the time of the study the most recent commercial model from Anthropic with the strongest performance among evaluated systems—with adopting the personas of real human panelists from the Center’s ongoing American Trends Panel (ATP).
The AI model was fed an extensive dossier of information for each panelist, including self-reported demographic data and detailed responses from Pew’s 2025 political typology survey. The model was then presented with the exact questions, order, instructions, and programming used in three distinct ATP survey waves administered during the first half of 2026.
Key findings from the study include:
- The Flattening of Diversity: AI respondents consistently eliminated the wide-ranging opinions found in human populations, gravitating heavily toward middle-of-the-road categories.
- Massive Discrepancies on Specific Issues: For instance, the AI model estimated that 98% of Americans believe China benefits more from U.S.-China trade relations—a figure 56 percentage points higher than the actual share of real U.S. adults.
- Zero-Selection Anomalies: While it is exceptionally rare for a human-answered survey option to receive zero responses, nearly half (47%) of the questions in the AI-generated synthetic poll featured at least one answer choice that was completely ignored by the model.
Despite publishing these findings, Pew Research Center emphasized that it has no current or future plans to utilize AI models to generate public opinion survey results, maintaining that direct engagement with the public remains essential to measuring authentic sentiment.

Chronology of the Experiment
The research project was structured carefully across late 2025 and mid-2026 to ensure a rigorous comparison between human benchmarks and AI performance.
- 2025 (The Foundation): Pew Research Center administered its comprehensive political typology survey to members of the American Trends Panel (ATP), gathering baseline ideological, demographic, and behavioral profiles that would later serve as the foundational "conditioning information" for the AI personas.
- First Half of 2026 (Human Survey Waves): Pew conducted three standard survey waves among human panelists on the ATP:
- Wave 185: Administered in early 2026, focusing heavily on views of political figures, including Donald Trump (questionnaires deployed in January).
- Wave 190: Administered in late April 2026, focusing on the United States’ role in the world and international trade.
- Wave 192: Administered in May 2026, measuring public perceptions of national problems.
- April 20–26, 2026: The original human panel data for the targeted waves was finalized. To make an exact, direct comparison, Pew isolated a subset of ATP panelists who had completed both the 2025 political typology survey and the relevant 2026 waves.
- April 27 – May 1, 2026 (The AI Replication): Researchers deployed Anthropic’s Claude Opus 4.6, configured with low reasoning settings, extended profile information, and expert reflection. The model was commanded to step into the shoes of the digital twins and answer the questionnaires sequentially.
- September 30, 2026: Pew Research Center officially published its comprehensive report, titled "Can AI Stand In for Human Survey-Takers? Not Really," alongside granular data tables, methodology notes, and questionnaire documentation.
Supporting Data and Analysis
The empirical data gathered during the experiment highlights stark differences between how real humans and synthetic AI respondents answer questions spanning geopolitics, personal behavior, and political attitudes.
Geopolitics and Trade
When asked whether the United States or China benefits more from bilateral trade, the divergence between human and machine was staggering. Real human respondents demonstrated a spread of opinions: 42% said China benefits more, 10% said the U.S. benefits more, 24% said both benefit equally, and 19% remained unsure.
In contrast, the synthetic AI poll estimated that an overwhelming 98% of Americans believed China benefited more, effectively erasing the nuance, uncertainty, and minority viewpoints present in the human population.
Personal Habits and Reported Behaviors
The AI’s tendency to compress responses into middle-tier categories was equally pronounced when evaluating daily lifestyle habits, such as sleep patterns and travel.
- Hours of Sleep: When asked how many hours of sleep they get per night, 54% of real U.S. adults reported getting 7 to 9 hours, while 39% reported 5 to 6 hours. Only 5% reported 4 hours or less, and 2% reported 10 hours or more. The AI synthetic poll dramatically distorted these numbers, claiming that 69% got 5 to 6 hours and 31% got 7 to 9 hours—while estimating that zero Americans fell into the extreme low (≤4 hours) or high (≥10 hours) categories, despite nearly one-in-ten real humans falling into these brackets.
- Trouble Sleeping: On a question measuring frequency of sleep disruption, 98% of synthetic respondents clustered into safe middle answers ("some days" or "rarely"), completely missing the 30% of real Americans who experience trouble every day, most days, or never.
- International Travel: While the AI model roughly approximated general travel trends, it estimated that only 1% of Americans had visited 10 or more countries outside the United States. In reality, 13% of U.S. adults reported reaching that milestone.
Political Evaluations: The Donald Trump Descriptors
In evaluating personal characteristics of political figures, the synthetic model routinely shied away from the extremes of measurement scales (such as "very well" or "not at all well"), instead favoring moderate options like "fairly well" or "not too well."
For example, when asked if phrases like "a good role model" or "keeps his promises" described Donald Trump:

- A Good Role Model: While 9% of real U.S. adults said this described him "very well," 0% of synthetic respondents chose that option.
- Keeps His Promises: While 16% of real Americans said he keeps his promises "very well" and 40% said "not at all well," the AI model assigned just 5% to the "very well" category and a mere 3% to "not at all well," instead lumping 92% of responses into the middle categories ("fairly well" or "not too well").
Conversely, the model did not always moderate its stance. When evaluating a list of 13 major national problems, the AI model occasionally over-indexed on extreme options, labeling several issues as "very big problems" at rates significantly higher than human panelists did.
Official Responses and Industry Context
The rise of "synthetic samples"—using Large Language Models to simulate focus groups and polling cohorts—has been a subject of intense debate within the data science and market research communities throughout 2026. Proponents argue that AI polling offers a faster, cheaper alternative to traditional human recruitment, which is becoming increasingly expensive and difficult.
However, Pew’s research provides a strong cautionary note against substituting human voices with silicon simulations. In its published policy statements, the Center drew a firm line: "The Center believes that speaking to the public is essential to measuring public opinion and has no current or future plans to use AI models to generate survey results."
The experiment underscores that while AI models can ingest vast amounts of demographic conditioning and mirror baseline trends, they lack authentic human consciousness, lived unpredictability, and the chaotic diversity of real public sentiment.
Implications for the Future of Polling
The implications of Pew’s study extend far beyond academic curiosity, striking at the heart of modern data collection, journalism, and democratic representation.
- The Danger of Echo Chambers and Homogeneity: If political campaigns, corporations, or media outlets begin relying on synthetic polling to gauge public mood, they risk receiving a sanitized, homogenized version of reality. The AI’s systemic erasure of minority viewpoints and outliers could blind decision-makers to fringe movements that eventually bubble up into mainstream political forces.
- Methodological Guardrails: The finding that 47% of survey questions featured "zero-selection" options in the AI poll demonstrates that foundational language models carry inherent biases toward or away from certain linguistic constructions and rating scales. Researchers who attempt to use AI as a shortcut must account for these artificial skews.
- The Irreplaceable Value of Human Contact: Ultimately, Pew’s research reinforces the premise that public opinion polling is fundamentally a social bridge connecting institutions to the lived experiences of citizens. When human voices are replaced by digital twins, the resulting data risks telling a story about what an algorithm thinks people feel, rather than what people actually experience.
