EconPapers    
Economics at your fingertips  
 

READ THE DAMN DOCUMENTATION (CAREFULLY). A case study using the PISA data

John Jerrim, Maria Palma Carvajal, Jake David Anders, María Ladrón de Guevara Rodríguez and Oscar David Marcenaro-Gutierrez
Additional contact information
Jake David Anders: UCL

No px84g_v1, SocArXiv from Center for Open Science

Abstract: When you get access to a new dataset, do you always carefully read the documentation first? We all know we should. But – let’s be honest – it’s a lot more fun to just start playing with the data. This can however be a dangerous game to play. This paper presents a case study of this matter using the OECD’s Programme for International Student Assessment (PISA). A survey question included in this study attempts to measure student truancy across countries over time. The international survey documentation suggests an identical question has been used across countries and cycles. Yet the national documentation illustrates how a subtle – yet important – change to the wording was made in some countries in 2015. We demonstrate how researchers could easily miss this change and how this would impact inferences in changes in truancy rates before and after the COVID-19 pandemic. Attempts to use artificial intelligence and large language models to spot this problem resulted in overconfidently incorrect advice. The findings thus serve as a reminder to even the most experienced data analysts (including ourselves) – ALWAYS READ THE SURVEY DOCUMENTATION CAREFULLY.

Date: 2025-11-11
References: Add references at CitEc
Citations:

Downloads: (external link)
https://osf.io/download/69134618d6bcbe4274e2f01e/

Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.

Export reference: BibTeX RIS (EndNote, ProCite, RefMan) HTML/Text

Persistent link: https://EconPapers.repec.org/RePEc:osf:socarx:px84g_v1

DOI: 10.31219/osf.io/px84g_v1

Access Statistics for this paper

More papers in SocArXiv from Center for Open Science
Bibliographic data for series maintained by OSF ().

 
Page updated 2025-11-16
Handle: RePEc:osf:socarx:px84g_v1