mlim: Single and multiple imputation with automated machine learning
E.F. Haghish
Additional contact information
E.F. Haghish: University of Bergen
2026 Stata Conference from Stata Users Group
Abstract:
Missing data are common in empirical research. Standard imputation tools require users to specify model forms and tuning parameters before it is clear which model best fits each incomplete variable. As data become more complex, for example because of nested structures, mixed variable types, or low-prevalence categories, imputation becomes more difficult. mlim is a new Stata package that uses automated machine learning for single and multiple imputation in mixed datasets. For each incomplete variable, mlim builds and tunes a separate prediction model, rather than applying one predefined model to all variables. The default method is elastic net, with optional random forest, gradient boosting, and stacked ensemble models for more computationally demanding applications. The package supports large datasets with continuous, binary, multinomial, and ordinal variables and includes automatic balancing procedures to reduce bias when categorical variables contain rare levels. This presentation introduces mlim to the Stata community as a competent open-source software for single and multiple imputation. It discusses the motivation, workflow, strengths, limitations, and examples of how Stata users can apply mlim to impute missing data.
Date: 2026-10-03
References: Add references at CitEc
Citations:
Downloads: (external link)
http://repec.org/usug2026/US26_Haghish.pptx
Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.
Export reference: BibTeX
RIS (EndNote, ProCite, RefMan)
HTML/Text
Persistent link: https://EconPapers.repec.org/RePEc:boc:usug26:24
Access Statistics for this paper
More papers in 2026 Stata Conference from Stata Users Group Contact information at EDIRC.
Bibliographic data for series maintained by Christopher F Baum ().