Preprint
Predicting Radiologist Expertise from 3D Gaze Patterns During CT Interpretation
arXiv
arXiv
08/24/2026
DOI: 10.48550/arxiv.2608.23836
Abstract
Accurate interpretation of volumetric CT requires efficient navigation of 3D image volumes and attention to diagnostically relevant regions. While eye-tracking has been widely studied in 2D medical imaging, its use for expertise assessment in CT settings remains limited. We propose a gaze-informed transformer framework for expertise classification in thoracic CT. Using a DINOv2 backbone, radiologist fixation patterns are integrated into volumetric feature learning through (1) a learnable log-space bias in self-attention and (2) gaze-weighted pooling of patch embeddings. We trained and evaluated our approach on 182 CT reading sessions from five radiologists with varying levels of experience. On a held-out test set, the model achieves an ROC-AUC of 0.91 and F1 score of 0.86, outperforming adapted methods. These findings suggest that incorporating visual search behavior into transformers may support objective, process-based expertise assessment in radiology. Code is available via https://github.com/leiluk1/GazeToSkill.
Details
- Title: Subtitle
- Predicting Radiologist Expertise from 3D Gaze Patterns During CT Interpretation
- Creators
- Leila KhaertdinovaAnna AnikinaClaudia Mello-ThomsBulat Ibragimov
- Resource Type
- Preprint
- Publication Details
- arXiv
- DOI
- 10.48550/arxiv.2608.23836
- ISSN
- 2331-8422
- Publisher
- arXiv
- Language
- English
- Date posted
- 08/24/2026
- Academic Unit
- Roy J. Carver Department of Biomedical Engineering; Radiology
- Record Identifier
- 9985220944202771
Metrics
1 Record Views