Anonymization and Information Loss
Ke Wu,
Baozhong Yang,
Zhenkun Ying and
Dexin Zhou
Papers from arXiv.org
Abstract:
Anonymizing financial texts prevents large language models (LLMs) from exploiting look-ahead bias, but inadvertently weakens the extracted signal. We propose a framework disentangling this information loss from bias removal. Predicting S&P credit downgrades, raw texts significantly outperform anonymized texts. Look-ahead bias does not explain this difference; rather, anonymization degrades predictive performance by masking informative numerical and agent entities, fundamentally altering how LLMs interpret the remaining context. This degradation spans various LLMs, tasks, and text types. To quantify it, we introduce the "anonymization gap," a metric requiring no outcome data.
Date: 2025-11, Revised 2026-09
References: View references in EconPapers View complete reference list from CitEc
Citations:
Downloads: (external link)
https://arxiv.org/pdf/2511.15364 Latest version (application/pdf)
Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.
Export reference: BibTeX
RIS (EndNote, ProCite, RefMan)
HTML/Text
Persistent link: https://EconPapers.repec.org/RePEc:arx:papers:2511.15364
Access Statistics for this paper
More papers in Papers from arXiv.org
Bibliographic data for series maintained by arXiv administrators ().