Reproducibility is not construct validity: LLM measurement of institutionally situated communication
Reproductibilité ne signifie pas validité de construit: l’usage des LLM pour mesurer des communications inscrites dans des contextes institutionnels
Veronika Batzdorfer () and
Carlo Romano Marcello Alessandro Santagiustina ()
Additional contact information
Veronika Batzdorfer: KIT - Karlsruhe Institute of Technology = Karlsruher Institut für Technologie
Carlo Romano Marcello Alessandro Santagiustina: ALMAnaCH - Automatic Language Modelling and ANAlysis & Computational Humanities - Centre Inria de Paris - Inria - Institut National de Recherche en Informatique et en Automatique, médialab - médialab (Sciences Po) - Sciences Po - Sciences Po, Sciences Po - Sciences Po
Working Papers from HAL
Abstract:
High annotation reproducibility does not necessarily imply that an LLM-inferred measure captures the construct it is intended to measure. We test this distinction using a dataset from the European Commission's AI Act consultation, linking structured survey responses to free-text consultation submissions from the same stakeholders. LLM annotations of consultation submissions are highly reproducible (intraclass correlations > 0.99), yet show limited convergence with survey-reported measures of the nominal construct they were intended to approximate. Divergence between survey-and LLM-inferred text-based measures varies systematically across stakeholder groups: business associations express greater concern about AI risks in text-based consultations than in survey responses (ḡ = +1.0), whereas public authorities and several nonbusiness groups show smaller or negative divergences. Divergences between scores suggest positive spatial autocorrelation across European countries (Moran's I = 0.347, p = 0.036), indicating that stakeholders from neighboring countries tend toward more similar text-based stances towards AI safety concerns. Despite divergence, survey-reported concerns remain strongly associated with support for explainability across all divergence levels. These results demonstrate that LLM annotation reproducibility can coexist with poor construct correspondence and motivate validation procedures that distinguish reproducibility, construct validity, and communication context variation when LLMs are used as measurement instruments.
Keywords: cross-national variation; AI regulation; culture and risk communication; risk perception; risk communication; construct validity; LLM annotation (search for similar items in EconPapers)
Date: 2026-08-31
Note: View the original document on HAL open archive server: https://hal.science/hal-05745153v1
References: Add references at CitEc
Citations:
Downloads: (external link)
https://hal.science/hal-05745153v1/document (application/pdf)
Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.
Export reference: BibTeX
RIS (EndNote, ProCite, RefMan)
HTML/Text
Persistent link: https://EconPapers.repec.org/RePEc:hal:wpaper:hal-05745153
Access Statistics for this paper
More papers in Working Papers from HAL
Bibliographic data for series maintained by CCSD ().