How economics classifies itself: text-based JEL codes and their consistency
Alessio Garau
MPRA Paper from University Library of Munich, Germany
Abstract:
Can a language model improve how economists classify their own papers? Only 15% of four million IDEAS/RePEc records carry usable JEL codes, and similar papers often receive different ones. I use a large language model (LLM) to solve this problem and assign three-digit codes from titles and abstracts, evaluating it on 69,503 coded articles published from 1991 to 2023 in the top 100 economics journals. Two tests do not assume that author codes provide the correct classification. Across semantic neighbors identified by a separate embedding model, model codes are 1.8 times as consistent as author codes. Holding codes per paper fixed, a blind check finds that 83% of model codes fit official American Economic Association (AEA) guidelines, compared with 67% of author codes. The classifier expands coverage, and the evaluation framework applies whenever human labels are incomplete or noisy.
Keywords: JEL codes; field classification; generative artificial intelligence; large language models; text as data; semantic similarity (search for similar items in EconPapers)
JEL-codes: A14 C45 C81 (search for similar items in EconPapers)
Date: 2026-07
References: Add references at CitEc
Citations:
Downloads: (external link)
https://mpra.ub.uni-muenchen.de/130163/1/MPRA_paper_130163.pdf original version (application/pdf)
Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.
Export reference: BibTeX
RIS (EndNote, ProCite, RefMan)
HTML/Text
Persistent link: https://EconPapers.repec.org/RePEc:pra:mprapa:130163
Access Statistics for this paper
More papers in MPRA Paper from University Library of Munich, Germany Ludwigstraße 33, D-80539 Munich, Germany. Contact information at EDIRC.
Bibliographic data for series maintained by Joachim Winter ().