An innovative feature selection method for support vector machines and its test on the estimation of the credit risk of default
Eduard Sariev and
Guido Germano
Review of Financial Economics, 2019, vol. 37, issue 3, 404-427
Abstract:
Support vector machines (SVM) have been extensively used for classification problems in many areas such as gene, text and image recognition. However, SVM have been rarely used to estimate the probability of default (PD) in credit risk. In this paper, we advocate the application of SVM, rather than the popular logistic regression (LR) method, for the estimation of both corporate and retail PD. Our results indicate that most of the time SVM outperforms LR in terms of classification accuracy for the corporate and retail segments. We propose a new wrapper feature selection based on maximizing the distance of the support vectors from the separating hyperplane and apply it to identify the main PD drivers. We used three datasets to test the PD estimation, containing (1) retail obligors from Germany, (2) corporate obligors from Eastern Europe, and (3) corporate obligors from Poland. Total assets, total liabilities, and sales are identified as frequent default drivers for the corporate datasets, whereas current account status and duration of the current account are frequent default drivers for the retail dataset.
Date: 2019
References: View references in EconPapers View complete reference list from CitEc
Citations:
Downloads: (external link)
https://doi.org/10.1002/rfe.1049
Related works:
Working Paper: An innovative feature selection method for support vector machines and its test on the estimation of the credit risk of default (2018) 
This item may be available elsewhere in EconPapers: Search for items with the same title.
Export reference: BibTeX
RIS (EndNote, ProCite, RefMan)
HTML/Text
Persistent link: https://EconPapers.repec.org/RePEc:wly:revfec:v:37:y:2019:i:3:p:404-427
Access Statistics for this article
More articles in Review of Financial Economics from John Wiley & Sons
Bibliographic data for series maintained by Wiley Content Delivery ().