Clustering for Fraud Detection
Abdelrahim Al Aqqad
Chapter Chapter 12 in Fraud Analytics in Action, 2026, pp 271-311 from Springer
Abstract:
Abstract This chapter introduces clustering as a foundational unsupervised machine learning technique for fraud detection in financial and transactional datasets. It begins by establishing the importance of defining normal behavior through exploratory data analysis and customer segmentation, recognizing that fraudulent activity is most effectively identified against a backdrop of well-characterized baseline patterns. The chapter then provides an in-depth treatment of the K-means algorithm, covering its iterative logic, sensitivity to centroid initialization, and the use of the elbow method to determine the optimal number of clusters. Practical Python implementations using scikit-learn—including MiniBatch K-means for scalability—demonstrate how to scale data, fit clustering models, and flag potential fraud cases by measuring each observation’s distance to its assigned cluster centroid. A comparative analysis with DBSCAN highlights the trade-offs between recall and precision, with DBSCAN achieving perfect recall at the cost of a higher false-positive rate. The chapter further explores alternative approaches including hierarchical clustering, self-organizing maps, principal component analysis, and autoencoders, equipping readers with a comprehensive toolkit for anomaly-based fraud detection in real-world settings.
Date: 2026
References: Add references at CitEc
Citations:
There are no downloads for this item, see the EconPapers FAQ for hints about obtaining it.
Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.
Export reference: BibTeX
RIS (EndNote, ProCite, RefMan)
HTML/Text
Persistent link: https://EconPapers.repec.org/RePEc:spr:sprchp:978-3-032-16023-2_12
Ordering information: This item can be ordered from
http://www.springer.com/9783032160232
DOI: 10.1007/978-3-032-16023-2_12
Access Statistics for this chapter
More chapters in Springer Books from Springer
Bibliographic data for series maintained by Sonal Shukla () and Springer Nature Abstracting and Indexing ().