Revealing economic facts: LLMs know more than they say
Marcus Buckmann,
Quynh Anh Nguyen and
Ed Hill
Additional contact information
Marcus Buckmann: Bank of England
Quynh Anh Nguyen: Bank of England
Ed Hill: Bank of England
No 1150, Bank of England Staff Working Paper series from Bank of England
Abstract:
We investigate whether hidden states of large language models (LLMs) can be used to estimate and impute economic and financial statistics. Focusing on county-level (eg unemployment) and firm-level (eg total assets) variables, we show that a linear regression trained on the hidden states of open-source LLMs outperforms the models' own text outputs. This indicates that internal representations encode richer economic information than is revealed directly in generated responses. A learning curve analysis shows that, in many cases, only a few dozen labelled examples suffice for training. We further propose a transfer learning method that improves estimation accuracy without requiring any labelled data for the target variable. Finally, we demonstrate the practical utility of hidden states in data imputation and super-resolution tasks.
Keywords: Large language models; embeddings; economic statistics; data imputation. (search for similar items in EconPapers)
JEL-codes: C21 C45 C81 (search for similar items in EconPapers)
Pages: 36
Date: 2025-10-31
References: Add references at CitEc
Citations:
Downloads: (external link)
https://www.bankofengland.co.uk/-/media/boe/files/ ... re-than-they-say.pdf
Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.
Export reference: BibTeX
RIS (EndNote, ProCite, RefMan)
HTML/Text
Persistent link: https://EconPapers.repec.org/RePEc:boe:boeewp:023272
Access Statistics for this paper
More papers in Bank of England Staff Working Paper series from Bank of England Bank of England, Threadneedle Street, London, EC2R 8AH. Contact information at EDIRC.
Bibliographic data for series maintained by Research ().