EconPapers    
Economics at your fingertips  
 

Anonymization and Information Loss

Ke Wu, Baozhong Yang, Zhenkun Ying and Dexin Zhou

Papers from arXiv.org

Abstract: Anonymizing financial texts prevents large language models (LLMs) from exploiting look-ahead bias, but inadvertently weakens the extracted signal. We propose a framework disentangling this information loss from bias removal. Predicting S&P credit downgrades, raw texts significantly outperform anonymized texts. Look-ahead bias does not explain this difference; rather, anonymization degrades predictive performance by masking informative numerical and agent entities, fundamentally altering how LLMs interpret the remaining context. This degradation spans various LLMs, tasks, and text types. To quantify it, we introduce the "anonymization gap," a metric requiring no outcome data.

Date: 2025-11, Revised 2026-09
References: View references in EconPapers View complete reference list from CitEc
Citations:

Downloads: (external link)
https://arxiv.org/pdf/2511.15364 Latest version (application/pdf)

Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.

Export reference: BibTeX RIS (EndNote, ProCite, RefMan) HTML/Text

Persistent link: https://EconPapers.repec.org/RePEc:arx:papers:2511.15364

Access Statistics for this paper

More papers in Papers from arXiv.org
Bibliographic data for series maintained by arXiv administrators ().

 
Page updated 2026-09-10
Handle: RePEc:arx:papers:2511.15364