Self-reported binary gender prediction from personality traits: Alignment between machine learning importance and classical effect sizes
Heeseung Cho,
Yiyu Chen and
Christian Wallraven
PLOS ONE, 2026, vol. 21, issue 8, 1-16
Abstract:
Gender differences in personality have been characterized using effect sizes such as Cohen’s d and multivariate measures such as the Mahalanobis distance. However, as machine learning models are increasingly applied to personality classification, it remains unclear how model-based feature importance relates to these classical measures of group separation. This study examined this relationship using two large-scale personality datasets comprising 18,062 participants in the Big5 and 44,324 participants in the 16PF. We trained multiple linear and nonlinear machine learning models (Logistic Regression, Random Forest, Extra Trees, LightGBM, and Explainable Boosting Machine) to predict self-reported binary gender from trait scores and derived feature-importance rankings using SHAP and permutation importance. These rankings were compared with univariate effect sizes, standardized linear discriminant weights, and trait-level contributions to the Mahalanobis distance. Across 20 model–resampling configurations, trait-importance rankings were highly stable in both datasets. In the Big5, univariate effect size rankings closely matched model-based importance rankings, with rank correlations of 0.98 to 1.00. Conversely, in the 16PF, alignment between univariate effect size and model-based importance was moderate (ρ=0.56−0.64), whereas alignment with multivariate discriminant structure remained strong (ρ=0.83−0.85). These findings indicate that divergence between univariate and model-based importance is consistent with covariance-mediated multivariate weighting rather than the emergence of novel predictive structure. When trait intercorrelations are limited, univariate summaries approximate multivariate importance; when covariance is substantial, multivariate geometry reshapes trait contributions, and model-based explainability metrics closely track the resulting discriminant structure.
Date: 2026
References: Add references at CitEc
Citations:
Downloads: (external link)
https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0355193 (text/html)
https://journals.plos.org/plosone/article/file?id= ... 55193&type=printable (application/pdf)
Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.
Export reference: BibTeX
RIS (EndNote, ProCite, RefMan)
HTML/Text
Persistent link: https://EconPapers.repec.org/RePEc:plo:pone00:0355193
DOI: 10.1371/journal.pone.0355193
Access Statistics for this article
More articles in PLOS ONE from Public Library of Science
Bibliographic data for series maintained by plosone ().