Data Preparation Techniques
Abdelrahim Al Aqqad
Chapter Chapter 6 in Fraud Analytics in Action, 2026, pp 113-140 from Springer
Abstract:
Abstract Data preparation is a foundational step in fraud analytics, transforming raw, unstructured data into reliable, analysis-ready datasets. This chapter covers the essential techniques required to clean and prepare fraud data, including handling missing values, detecting and treating outliers, data integration, sampling, and discretization. Missing value management is examined through methods such as deletion, mean/mode imputation, regression imputation, and specialized approaches including LOCF, NOCB, and KNN imputation. Outlier detection is explored using statistical methods such as Z-score and interquartile range (IQR), as well as machine learning approaches including isolation forests, one-class SVM, and elliptic envelope, with treatment techniques covering truncation, capping, and winsorizing. The chapter further discusses data integration through record linkage and entity resolution, data sampling strategies such as simple random, stratified, and cluster sampling, and data discretization methods including uniform-width binning, equal frequency binning, and k-means clustering. Practical examples with real-world fraud datasets illustrate each technique, equipping data analysts with the skills to improve data quality and reliability as a prerequisite to building accurate and effective fraud detection models.
Date: 2026
References: Add references at CitEc
Citations:
There are no downloads for this item, see the EconPapers FAQ for hints about obtaining it.
Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.
Export reference: BibTeX
RIS (EndNote, ProCite, RefMan)
HTML/Text
Persistent link: https://EconPapers.repec.org/RePEc:spr:sprchp:978-3-032-16023-2_6
Ordering information: This item can be ordered from
http://www.springer.com/9783032160232
DOI: 10.1007/978-3-032-16023-2_6
Access Statistics for this chapter
More chapters in Springer Books from Springer
Bibliographic data for series maintained by Sonal Shukla () and Springer Nature Abstracting and Indexing ().