Logo image
Interpretable explanatory item response models with machine learning
Dissertation

Interpretable explanatory item response models with machine learning

Xiaoting Zhong
University of Iowa
Doctor of Philosophy (PhD), University of Iowa
Spring 2026
DOI: 10.25820/etd.008346
pdf
Thesis-XZ2.38 MB
Embargoed Access, Embargo ends: 06/29/2028

Abstract

Educational assessment increasingly requires models that are not only accurate, but also interpretable and fair across diverse groups of learners. Traditional psychometric models, including item response theory (IRT) and explanatory item response models (EIRMs), provide a principled framework for modeling person- and item-level variation, yet they often rely on restrictive assumptions that may fail to capture complex nonlinearities and interactions in modern educational data. In contrast, machine learning (ML) methods offer substantial predictive flexibility, but they typically ignore the hierarchical person-item structure central to psychometric modeling and are often criticized for limited transparency. This dissertation addresses this gap by proposing a hybrid framework, Explanatory Item Response Models with Machine Learning Integration (EIRM-ML), which augments the explanatory structure of EIRMS with ML-based predicted probabilities while preserving random person and item effects. The dissertation is organized around three central questions: whether EIRM-ML improves predictive performance and calibration relative to conventional EIRMS and standalone ML models; whether interpretable ML tools can deepen substantive understanding of person- and item-level effects; and whether the hybrid approach supports more equitable prediction across demographic subgroups. To answer these questions, the study develops the conceptual and methodological foundations of EIRM-ML, presents a two-stage estimation strategy based on cross-validated ML predictions embedded in a mixed-effects psychometric model, and evaluates the framework using both simulation and empirical evidence. A comprehensive simulation study examines the performance of EIRM-ML under linear, nonlinear, interaction, and differential item functioning scenarios. The results are designed to show that the proposed framework remains competitive when the linear model is correctly specified and provides clear gains in discrimination, calibration, and subgroup robustness when nonlinearities or interaction effects are present. An empirical application using the Open University Learning Analytics Dataset (OULAD) illustrates how the framework can be implemented in practice to model student-item response behavior using both learner characteristics and item features. The application also demonstrates how feature importance measures, partial dependence plots, accumulated local effects, and subgroup calibration diagnostics can be used to interpret predictive mechanisms and assess fairness. This dissertation contributes to educational measurement by offering a unified framework that bridges psychometric explanation and algorithmic prediction. By combining hierarchical modeling, predictive flexibility, and interpretability, EIRM-ML provides a practical and theoretically grounded approach for next-generation assessment analytics.

Details

Metrics

1 Record Views
Logo image