Examining the joint impact of missing data mechanisms and item parameter drift on the accuracy of item response theory-based test equating
Ayse Bilicioglu Gunes and
Serife Zeybekoglu Yesil
PLOS ONE, 2026, vol. 21, issue 7, 1-17
Abstract:
Maintaining score comparability across different test administrations is essential in large-scale educational and psychological assessment. Two major threats to equating accuracy are item parameter drift (IPD), reflecting changes in anchor item characteristics over time, and missing item responses, which commonly occur in operational testing. Although each factor has been studied separately, their joint impact on equating accuracy has not been systematically evaluated. A Monte Carlo simulation was conducted using a nonequivalent groups with anchor test design. Data were generated under a three-parameter logistic model with 2,000 examinees per form and 500 replications per condition. The Stocking–Lord method was used to estimate equating constants across 25 conditions, varying IPD rate (0%, 10%, 20%), IPD magnitude (0, 0.25, 0.50), missing data mechanism (MCAR vs. MAR), and missing data rate (0%, 10%, 20%). Missing responses were addressed using multiple imputation by chained equations. Equating accuracy was evaluated using bias and root mean square error. Results indicated that the B constant was highly sensitive to the missing data mechanism: MCAR conditions produced near-zero bias, whereas MAR conditions introduced substantial positive bias, particularly at higher missing rates. IPD alone did not result in meaningful bias but increased estimation variability. When IPD and MAR co-occurred, equating error exceeded the sum of their individual effects, indicating an interaction rather than a purely additive relationship. Although multiple imputation reduced error under MCAR, it did not fully eliminate bias under MAR conditions. In contrast, the scale constant remained stable across all conditions. Overall, the missing data mechanism had a stronger impact on equating accuracy than either the rate of missingness or the degree of item drift alone. When both factors were present, their combined influence led to increased error that could not be fully corrected through imputation. These findings underscore the importance of carefully evaluating missing data mechanisms, selecting appropriate imputation strategies, such as MICE, and monitoring anchor item stability to ensure accurate score equating, particularly in psychological assessment contexts where even small errors may affect individual-level decisions.
Date: 2026
References: Add references at CitEc
Citations:
Downloads: (external link)
https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0353665 (text/html)
https://journals.plos.org/plosone/article/file?id= ... 53665&type=printable (application/pdf)
Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.
Export reference: BibTeX
RIS (EndNote, ProCite, RefMan)
HTML/Text
Persistent link: https://EconPapers.repec.org/RePEc:plo:pone00:0353665
DOI: 10.1371/journal.pone.0353665
Access Statistics for this article
More articles in PLOS ONE from Public Library of Science
Bibliographic data for series maintained by plosone ().