LLM measurements that agree with human coders overall can still distort specific effects
Alexander K. Moore,
Bingqing Li,
Shane Wang and
Matt Thomson
No dz3qe_v1, SocArXiv from Center for Open Science
Abstract:
Large language models (LLMs) are increasingly adopted in the social sciences as surrogates for human coders. Standard practice validates them by their correspondence with human measurements of the same stimuli using statistics like correlation. However, overall correspondence can conceal biases that distort specific claims. We demonstrate this with open-weight LLM ratings of 568 food and 1,811 scene photographs. Ratings inflate, attenuate, or reverse effects, and improving overall correspondence with prompts can improve one contrast while worsening another. Rankings also elevate some categories and demote others beyond what random errors of the same size typically produce. Across six prompts, some of these errors are shared, so sampling across prompts may continue to mask biases. We propose validating LLM coders against specific claims they will test rather than the general coding task.
Date: 2026-09-21
References: Add references at CitEc
Citations:
Downloads: (external link)
https://osf.io/download/6aaef5aa5baea946b3ebb13b/
Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.
Export reference: BibTeX
RIS (EndNote, ProCite, RefMan)
HTML/Text
Persistent link: https://EconPapers.repec.org/RePEc:osf:socarx:dz3qe_v1
DOI: 10.31235/osf.io/dz3qe_v1
Access Statistics for this paper
More papers in SocArXiv from Center for Open Science
Bibliographic data for series maintained by OSF ().