Journal article
Online, Crowdsourced Sampling (OCS) Platforms for Large-Scale Data Collection in Voice Disorders Research: Prices, Pitfalls, and Perks
Journal of voice
08/25/2026
DOI: 10.1016/j.jvoice.2026.07.049
PMID: 42642244
Abstract
Online, crowdsourced sampling (OCS) platforms such as Amazon’s Mechanical Turk and CloudResearch Connect enable faster, cheaper, and larger-scale data collection than in-person sampling methods. However, OCS platforms are often criticized due to data quality concerns. The current study examines a battery of screening tools to determine which most effectively predict data quality and to quantify the influence of low-quality data on the statistical properties of the samples. Finally, we issue recommendations to researchers interested in OCS-based clinical voice research that may reduce costly missteps when fielding their first OCS studies.
To evaluate the suitability of OCS for clinical voice research, survey-based measures of vocal function, vocal fatigue, personality, communicative quality of life, and other factors were collected on Qualtrics from (1) a convenience sample of undergraduate students (n = 47), (2) Mechanical Turk users (n = 495), and (3) Connect users (n = 99) over six months. Low-quality responses were eliminated using six screening tools in four nested stages. Logistic regression was used to identify variables that strongly predicted data quality. Finally, patient-reported outcome measure (PROM) responses were compared between the high- and low-quality groups.
Depending on the stage, data quality screenings flagged between 5.00% (n = 32) and 13.6% (n = 87) of responses as low-quality. Hispanic/Latino ethnicity and long response times most strongly predicted low-quality responses, which uniformly exhibited more severe ratings of health- and voice-related disability than did high-quality responses.
Content-based screenings, particularly analysis of responses to open-ended questions, meaningfully augmented automated screenings. Low-quality respondents systematically over-reported health- and voice-related disability, a severity bias with plausible economic and behavioral explanations. The emergence of AI-generated responses represents an evolving and distinct challenge that content-based screening may be increasingly insufficient to address. We provide recommendations for collecting high-quality data on OCS platforms to minimize barriers to entry for future clinical voice research.
Details
- Title: Subtitle
- Online, Crowdsourced Sampling (OCS) Platforms for Large-Scale Data Collection in Voice Disorders Research: Prices, Pitfalls, and Perks
- Creators
- Christopher S. Apfelbach - University of MinnesotaLady Catherine Cantor-Cutiva - East Tennessee State UniversityEric J. Hunter - University of Iowa
- Resource Type
- Journal article
- Publication Details
- Journal of voice
- DOI
- 10.1016/j.jvoice.2026.07.049
- PMID
- 42642244
- ISSN
- 0892-1997
- eISSN
- 1873-4588
- Publisher
- Elsevier Inc
- Language
- English
- Electronic publication date
- 08/25/2026
- Academic Unit
- Communication Sciences and Disorders; Teaching and Learning; Otolaryngology
- Record Identifier
- 9985220286102771
Metrics
1 Record Views