Logo image
A comparative study of methodological approaches to vertical scale quality
Dissertation   Open access

A comparative study of methodological approaches to vertical scale quality

Guangyun Liu
University of Iowa
Doctor of Philosophy (PhD), University of Iowa
Spring 2026
DOI: 10.25820/etd.008371
pdf
A Comparative Study of Methodological Approaches to Vertical Scale Quality7.70 MBDownloadView
Open Access

Abstract

The primary function of vertical scaling is to measure growth. Grade-level students’ scores can be put onto the same vertical scale so they can be directly compared. The complexity of vertical scaling stems from the many decisions involved in the process, especially when using Item Response Theory (IRT)-based scaling methods. This dissertation addressed the need for a comprehensive study to examine the many factors contributing to a vertical scale and to explore potential interactions. The simulation study factors include data collection design, IRT model, calibration methods, linking methods, and examinee population characteristics. In addition, two examples are included with code and immediate interpretation of results to illustrate the procedure for conducting vertical scaling. This dissertation highlighted the role of linking methods and introduced simultaneous linking (SM) in a vertical scaling setting. Simultaneous linking is a regression-based method with great operational convenience, as it can link multiple test forms at the same time. SM’s performance was evaluated against Stocking-Lord’s transformation (SL) and fixed calibration (FE). All three linking methods are able to achieve comparable item parameter recovery, proficiency distribution recovery, and scale score recovery. Although the comparative performance of the linking methods differed across conditions, they are all able to produce very similar vertical scales. However, varying linking methods did not affect scale quality significantly, especially compared to higher-level decisions such as data collection design, IRT model, and calibration methods. The relative performance of the three linking methods was more stable under the two-parameter logistic (2-PL) model and common-item (CI) design. Under the one-parameter logistic (1-PL) model and scaling test (ST) design, the differences among linking methods were negligible, and all produced more accurate vertical scales. One key observation in this dissertation is that the performance of the three linking methods at the proficiency (θ) level does not always translate to the scale score level, even though only a linear transformation was applied. This raised the concern that how scaling is conducted might play an important role that was overlooked, as extreme values can distort the scale transformation. Overall, the results suggest that vertical scaling is a context-dependent procedure and is very sensitive to changes in higher-level factors. Besides certain uncontrollable factors like examinee population characteristics, practitioners should understand the pros and cons of each method to make an informed decision tailored to their scaling scenario.
Fixed Calibration IRT Linking Simultaneous Linking Stocking-Lord Vertical Scaling

Details

Metrics

1 Record Views
Logo image