Comparative Study of Supervised and Unsupervised Learning Models in Fraud Detection
Adaobi Beverly Akonobi and
Christiana Onyinyechi Okpokwu
International Journal of Scientific Research in Computer Science, Engineering and Information Technology, 2024, vol. 10, issue 4, 777-820
Abstract:
Fraud detection has become a critical concern for financial institutions and e-commerce platforms due to the increasing sophistication of fraudulent activities. Machine learning offers powerful tools for identifying anomalous patterns and fraudulent behavior in large, complex datasets. This study presents a comparative analysis of supervised and unsupervised learning models in the context of fraud detection, focusing on their methodologies, performance, and applicability. Supervised models, such as Logistic Regression, Decision Trees, Random Forest, and Support Vector Machines (SVM), require labeled datasets and learn to classify transactions as fraudulent or legitimate based on historical data. These models typically achieve high accuracy when trained on well-labeled, balanced datasets. However, they are limited by the availability and quality of labeled data, especially in environments where fraud evolves rapidly. Unsupervised models, including K-Means Clustering, Isolation Forest, and Autoencoders, do not require labeled data and are effective at detecting novel or emerging fraud patterns by identifying outliers and anomalies. While these models are more adaptable to unseen fraud types, they may produce higher false-positive rates due to the absence of explicit class labels. This paper compares the performance of both approaches using benchmark datasets and evaluates them based on precision, recall, F1-score, and Area Under the Curve (AUC). Results show that while supervised models perform better in stable environments with abundant labeled data, unsupervised models excel in detecting new fraud cases and offer greater flexibility in dynamic environments. Additionally, the paper explores hybrid and semi-supervised approaches that combine the strengths of both paradigms, aiming to enhance detection accuracy while reducing dependency on labeled data. The study concludes that the choice of model depends on data availability, operational constraints, and the evolving nature of fraud. A balanced strategy that incorporates both supervised and unsupervised techniques is recommended for robust and scalable fraud detection systems in real-world applications.
Keywords: Fraud Detection; Supervised Learning; Unsupervised Learning; Machine Learning; Anomaly Detection; Classification Models; Autoencoders; Outlier Detection; Financial Fraud; Hybrid Models (search for similar items in EconPapers)
Date: 2024
Note: Article URL: https://ijsrcseit.com/home/article/view/CSEIT25113496
References: Add references at CitEc
Citations:
Downloads: (external link)
https://ijsrcseit.com/home/article/view/CSEIT25113496 Article URL (text/html)
https://ijsrcseit.com/home/article/download/CSEIT25113496/CSEIT25113496 Full text (application/pdf)
Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.
Export reference: BibTeX
RIS (EndNote, ProCite, RefMan)
HTML/Text
Persistent link: https://EconPapers.repec.org/RePEc:jbh:ijsrcs:v10:y2024:i4:id:1645
DOI: 10.32628/CSEIT25113496
Access Statistics for this article
More articles in International Journal of Scientific Research in Computer Science, Engineering and Information Technology from International Journal of Scientific Research in Computer Science, Engineering and Information Technology
Bibliographic data for series maintained by Pankaj Sharma (USA) ().