Skip to main content
QUICK REVIEW

[Paper Review] An IRT-based Model for Omitted and Not-reached Items

Jinxin Guo|arXiv (Cornell University)|Apr 7, 2019
Psychometric Methodologies and TestingDecision Sciences31 references3 citations
TL;DR

This paper proposes a novel IRT-based model that distinguishes between omitted and not-reached items in educational and psychological testing by modeling nonignorable missingness through a cumulative missingness mechanism linked to examinee ability. Using Bayesian estimation via MCMC, the model demonstrates superior parameter recovery and model fit compared to listwise deletion and ignorable models, particularly in nonignorable missingness scenarios.

ABSTRACT

Missingness is a common occurrence in educational assessment and psychological measurement. It could not be casually ignored as it may threaten the validity of the test if not handled properly. Considering the difference between omitted and not-reached items, we developed an IRT-based model to handle these missingness. In the proposed method, not-reached responses are captured by the cumulative missingness. Moreover, the nonignorability is attributed to the correlation between ability and person missing trait. We proved that its item parameters estimate under maximum marginal likelihood (MML) estimation is consistent. We further proposed a Bayesian estimation procedure using MCMC methods to estimate all the parameters. The simulation results indicate that the model parameters under the proposed method are better recovered than that under listwise deletion, and the nonignorable model fits the simulated nonignorable nonresponses better than ignorable model in terms of Bayesian model selection. Furthermore, the Program for International Student Assessment (PISA) data set was analyzed to further illustrate the usage of the proposed method.

Motivation & Objective

  • Address the challenge of nonignorable missing data in educational and psychological assessments, where missing responses stem from both omission and non-reach due to time limits.
  • Differentiate between omitted items (skipped intentionally) and not-reached items (unreached due to time constraints), which are often conflated in existing models.
  • Develop a statistically consistent IRT-based model that accounts for the correlation between examinee ability and missingness propensity, ensuring valid parameter estimation.
  • Provide a Bayesian estimation framework using MCMC to jointly estimate IRT parameters and missing data mechanisms, improving robustness under nonignorable missingness.
  • Evaluate model performance using simulation studies and real-world PISA data to demonstrate superiority over traditional listwise deletion and ignorable missingness models.

Proposed method

  • Formalize the missing data mechanism using a nonignorable model where missingness depends on both observed responses and examinee ability, capturing the correlation between ability and missingness trait.
  • Model not-reached responses through a cumulative missingness process, distinguishing them from omitted items via a latent variable structure.
  • Construct a joint likelihood model based on selection models (SLM), factorizing the joint distribution of responses and missingness indicators as P(R,Y|Ω) = P(R|Y,Ω)P(Y|Ω).
  • Implement Bayesian estimation via MCMC with Gibbs sampling, using informative priors for IRT parameters (a_j, b_j, ζ_j) and missingness parameters (γ_0, γ_1, γ_2), with constraints to reflect theoretical expectations.
  • Use the deviance information criterion (DIC) and logarithm of the pseudomarginal likelihood (LPML) for Bayesian model comparison, with numerical stabilization via U_ij,max in CPO estimation.
  • Apply the model to both simulated data and real PISA data to validate performance and interpretability across diverse missingness patterns.

Experimental results

Research questions

  • RQ1How does the proposed IRT model perform in recovering item and ability parameters when missingness is nonignorable and stems from both omitted and not-reached items?
  • RQ2Can the model distinguish between omitted and not-reached items more effectively than existing methods that treat them as a single missingness type?
  • RQ3Does the Bayesian MCMC estimation procedure yield consistent and reliable parameter estimates under nonignorable missingness, particularly when listwise deletion leads to bias?
  • RQ4How does the proposed nonignorable model compare to ignorable models in terms of model fit using DIC and LPML in simulated and real data?
  • RQ5What is the empirical performance of the model when applied to real PISA data, particularly in capturing the relationship between examinee ability and missingness patterns?

Key findings

  • The proposed model achieved better parameter recovery for both IRT item parameters and missingness parameters compared to listwise deletion, especially under high nonignorable missingness.
  • Bayesian model selection using DIC and LPML showed that the nonignorable model provided a significantly better fit to simulated nonignorable nonresponses than the ignorable model.
  • The model's maximum marginal likelihood (MML) estimates were proven to be consistent, ensuring valid inference under the proposed missingness mechanism.
  • In the PISA data analysis, the model successfully captured the nonignorable nature of missingness, revealing that lower-ability students were more likely to omit or not reach items.
  • The MCMC sampling procedure demonstrated stable convergence and reliable posterior estimation, with effective sample sizes sufficient for inference.
  • The use of U_ij,max in CPO computation ensured numerical stability in LPML estimation, avoiding issues common in high-dimensional missing data models.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.