Handling Imbalanced Data Sets
Abdelrahim Al Aqqad
Chapter Chapter 9 in Fraud Analytics in Action, 2026, pp 201-232 from Springer
Abstract:
Abstract Imbalanced datasets—where one class significantly outnumbers another—present a fundamental challenge in machine learning, particularly in fraud analytics where fraudulent transactions constitute a small minority of all records. This chapter provides a comprehensive examination of class imbalance, its origins, and the practical consequences it imposes on model performance, including majority-class bias, elevated false-positive rates, and failure to detect the minority class. The chapter introduces evaluation metrics suited to imbalanced settings—precision, recall, F1-score, and AUC-ROC—and explains why conventional accuracy measures are inadequate in such contexts. A range of resampling and data augmentation strategies are examined in detail, including random under-sampling, random oversampling, and the Synthetic Minority Over-Sampling Technique (SMOTE) with its variants such as Borderline-SMOTE, ADASYN, Safe-Level SMOTE, DBSMOTE, and MWMOTE. The chapter also covers cost-sensitive learning, class weight balancing, ensemble methods, transfer learning, hybrid approaches, and the adjustment of posterior probability estimates. Practical examples grounded in credit card and insurance fraud scenarios, supported by hands-on Python laboratory exercises, illustrate each technique. Guidance on when and how to apply each method—including the critical principle that resampling must be applied only to training data—equips practitioners to build robust, unbiased fraud detection models.
Date: 2026
References: Add references at CitEc
Citations:
There are no downloads for this item, see the EconPapers FAQ for hints about obtaining it.
Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.
Export reference: BibTeX
RIS (EndNote, ProCite, RefMan)
HTML/Text
Persistent link: https://EconPapers.repec.org/RePEc:spr:sprchp:978-3-032-16023-2_9
Ordering information: This item can be ordered from
http://www.springer.com/9783032160232
DOI: 10.1007/978-3-032-16023-2_9
Access Statistics for this chapter
More chapters in Springer Books from Springer
Bibliographic data for series maintained by Sonal Shukla () and Springer Nature Abstracting and Indexing ().