Logo image
Fairness, validity, and impact of a standardized English proficiency test in EMI contexts and beyond: a mixed-methods study of test takers’ perceptions
Dissertation

Fairness, validity, and impact of a standardized English proficiency test in EMI contexts and beyond: a mixed-methods study of test takers’ perceptions

I-Chun Hsiao
University of Iowa
Doctor of Philosophy (PhD), University of Iowa
Spring 2026
DOI: 10.25820/etd.008406
pdf
Dissertation_I-Chun Hsiao2.84 MB
Embargoed Access, Embargo ends: 06/29/2028

Abstract

While English proficiency tests such as TOEFL, IELTS, and the Duolingo English Test have been extensively examined in assessment research, locally developed large-scale assessments remain underexplored, particularly in non-English-dominant higher education contexts. This study addresses this gap by investigating test takers’ perceptions of the BEST Test of English Proficiency (BESTEP), Taiwan’s first academically oriented English proficiency test introduced under the Bilingual 2030 policy. Specifically, the study examines the BESTEP’s fairness, validity, and its broader impact on learning experiences and future opportunities. Building on Green’s (2014, 2020) Framework for Effective Assessment Systems, this study proposes a Quality of Standardized Language Tests model that conceptualizes test quality as an interconnected system comprising fairness, validity, impact, reliability, and practicality. Fairness is positioned at the same conceptual level as validity, drawing on Wallace’s (2018) Language Assessment Fairness Model, while validity is framed in accordance with the Standards for Educational and Psychological Testing (AERA et al., 2014). Test impact is conceptualized through Shih’s (2007) Washback Model and related work (Bachman & Palmer, 1996; ILTA, 2024), involving both intended and unintended consequences at individual and systemic levels. An explanatory sequential mixed-methods design was employed. Quantitative data were collected through an online survey (N = 424) adapted from Wallace and Qin (2021), Dong (2022), and Nguyen (2023). Content validity was established through expert review (scale-CVI = .93), and the final 48-item instrument demonstrated strong internal consistency (Cronbach’s α = .88–.93) and construct validity (composite reliability > .75; average variance extracted > .50). Qualitative data were gathered through semi-structured interviews with 15 test takers representing diverse institutional and disciplinary backgrounds. Quantitative analyses included descriptive statistics, confirmatory factor analysis, and nonparametric group comparisons, while qualitative data were thematically analyzed and member-checked to enhance credibility. Quantitative results indicated generally positive perceptions of fairness, relevance, and impact, though distributive fairness and perceived test utility received comparatively lower ratings. No significant differences emerged by gender, test-taking frequency, participation in English as a medium of instruction (EMI) courses, or year of study; however, institutional type and Ministry of Education (MOE) funding category were associated with significant differences in perceptions. Qualitative findings further revealed that fairness judgments extended beyond test content and scoring to involve administrative policies, access conditions, and interactional experiences during test administration. While BESTEP was viewed as broadly aligned with general language competence, concerns were raised regarding its alignment with EMI demands, its discriminatory power, and its difficulty level. In terms of impact, the test exerted limited influence on learners’ motivation, confidence, and study behaviors, and its perceived benefits were largely short-term and immediate. Participants also questioned its recognition for further study and employment, particularly in comparison with internationally established assessments. By centering test takers’ perspectives, this study advances a more context-sensitive and justice-oriented understanding of standardized language assessment. It demonstrates that test quality must be evaluated not only through psychometric rigor but also through stakeholder experiences and institutional contexts in which tests are implemented. The findings provide methodological guidance for mixed-methods research and for the adaptation and validation of perception-based survey instruments. They also provide practical and policy-relevant insights for locally developed large-scale assessments in Taiwan and other non-English-dominant higher education systems. 
English as a medium of instruction (EMI) Language assessment Test fairness Test impact Test validity Test-taker perceptions Foreign language education

Details

Metrics

1 Record Views
Logo image