Abstract
Empirical conclusions depend not only on data but on analytic decisions made throughout the research process. Many-analyst studies have quantified this: independent teams testing the same hypothesis on the same dataset often reach conflicting conclusions. But such studies require costly coordination and are rarely conducted. We show that fully autonomous AI analysts built on large language models (LLMs) can cheaply and at scale replicate this structured analytic diversity. In our framework, each AI analyst executes a complete analysis pipeline on a fixed dataset and hypothesis, while a separate AI auditor screens runs for methodological validity. Across three datasets, AI-generated analyses exhibit substantial dispersion in effect sizes, p-values, and conclusions, driven by systematic differences in preprocessing, model specification, and inference across LLMs and personas. Critically, outcomes are steerable: changing the analyst persona or model shifts the distribution of results even among valid analyses.
These findings highlight a central challenge for AI-automated empirical science: when defensible analyses are cheap, evidence becomes abundant and vulnerable to selective reporting. But the same capability suggests a solution: treating results as distributions makes analytic uncertainty visible, and deploying AI analysts on a fixed specification can reveal disagreement from underspecified choices. We therefore argue for new transparency norms: multiverse-style reporting and prompt disclosure, alongside code and data.
Joint work with Martin Bertran and Riccardo Fogliato