Randomness In Large Language Models: What Researchers Need to Know (And Report)
Guillaume Coqueret (),
Joan Llull,
Florian Oswald (),
Christophe Pérignon,
Christoph Scheuch and
Lars Vilhuber
Additional contact information
Guillaume Coqueret: EM - EMLyon Business School
Joan Llull: CSIC - Consejo Superior de Investigaciones Cientificas [España] = Spanish National Research Council [Spain], Barcelona School of Economics (Spain, Barcelona) - BSE
Florian Oswald: UNITO - Università degli studi di Torino = University of Turin
Christophe Pérignon: HEC Paris - Ecole des Hautes Etudes Commerciales
Christoph Scheuch: HU Berlin - Humboldt-Universität zu Berlin = Humboldt University of Berlin = Université Humboldt de Berlin
Lars Vilhuber: CU - Cornell University [Ithaca, NY, USA]
Working Papers from HAL
Abstract:
Large language models (LLMs) are increasingly used to generate data for research. Typical use cases are classifications, annotations, information extraction, and generation of numerical scores. Unlike conventional measurements, LLM outputs can vary across repeated requests even when the prompt and apparent model settings remain unchanged. This variation arises from deliberate sampling, silent model updates, numerical rounding, or expert routing. Setting a dedicated temperature parameter to zero removes deliberate sampling when that option is available, but it does not eliminate the other sources of randomness. Exact reproduction is therefore generally not possible when using proprietary application programming interfaces. Local execution of open-weight models offers greater control, but reproducibility still depends on the complete hardware and software stack. We illustrate these issues through sentiment classifications of corporate filings and examine their consequences for downstream regression results. We then propose a reporting standard for articles and replication packages, as well as guidance for data editors and authors. Together, these findings and recommendations establish that LLM outputs should be treated as draws from a distribution rather than as fixed measurements.
Keywords: Measurement Error; Text-as-Data; Reproducibility; Large Language Models (search for similar items in EconPapers)
Date: 2026-07-29
References: Add references at CitEc
Citations:
There are no downloads for this item, see the EconPapers FAQ for hints about obtaining it.
Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.
Export reference: BibTeX
RIS (EndNote, ProCite, RefMan)
HTML/Text
Persistent link: https://EconPapers.repec.org/RePEc:hal:wpaper:hal-05730402
DOI: 10.2139/ssrn.7191580
Access Statistics for this paper
More papers in Working Papers from HAL
Bibliographic data for series maintained by CCSD ().