EconPapers    
Economics at your fingertips  
 

Variable selection for categorical response: a comparative study

Sweata Sen, Damitri Kundu and Kiranmoy Das ()
Additional contact information
Sweata Sen: Indian Statistical Institute
Damitri Kundu: Indian Statistical Institute
Kiranmoy Das: Indian Statistical Institute

Computational Statistics, 2023, vol. 38, issue 2, No 10, 809-826

Abstract: Abstract Variable selection is a well-studied problem in linear regression, but the existing works mostly deal with continuous responses. However, in many applications, we come across data with categorical responses. In the classical (frequentist) approach there exists penalized regression methods (e.g. logistic Lasso) which can be used for variable selection when we have a categorical response, and a large number of predictors. In this paper, we compare the performance of three alternative approaches for handling data with a single categorical response and multiple continuous (or count) predictors. In addition to the well-known logistic Lasso, we consider a model-based Bayesian approach, and a model-free approach for variable selection. We consider a binary response, and a response with three categories. Through extensive simulation studies we compare the performance of these three competing methods. We observe that the model-based methods can often accurately identify the important predictors, but sometimes fail to detect the unimportant ones. Also the model-based approaches are computationally expensive whereas the model-free approach is extremely fast. For misspecified models, the model-free method really outperforms in prediction. However, when the predictors are correlated (moderately or substantially) then the model-based methods perform better than the model-free method. We analyse the well-known Pima Indian Diabetes dataset for illustrating the effectiveness of three competing methods under consideration.

Keywords: Bayesian Lasso; Categorical response; Data-augmentation; Gibbs sampler; Kolmogorov–Smirnov test; Logistic Lasso; MCMC (search for similar items in EconPapers)
Date: 2023
References: View references in EconPapers View complete reference list from CitEc
Citations:

Downloads: (external link)
http://link.springer.com/10.1007/s00180-022-01260-1 Abstract (text/html)
Access to the full text of the articles in this series is restricted.

Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.

Export reference: BibTeX RIS (EndNote, ProCite, RefMan) HTML/Text

Persistent link: https://EconPapers.repec.org/RePEc:spr:compst:v:38:y:2023:i:2:d:10.1007_s00180-022-01260-1

Ordering information: This journal article can be ordered from
http://www.springer.com/statistics/journal/180/PS2

DOI: 10.1007/s00180-022-01260-1

Access Statistics for this article

Computational Statistics is currently edited by Wataru Sakamoto, Ricardo Cao and Jürgen Symanzik

More articles in Computational Statistics from Springer
Bibliographic data for series maintained by Sonal Shukla () and Springer Nature Abstracting and Indexing ().

 
Page updated 2025-03-20
Handle: RePEc:spr:compst:v:38:y:2023:i:2:d:10.1007_s00180-022-01260-1