Logo image
Predicting Radiologist Expertise from 3D Gaze Patterns During CT Interpretation
Preprint   Open access

Predicting Radiologist Expertise from 3D Gaze Patterns During CT Interpretation

Leila Khaertdinova, Anna Anikina, Claudia Mello-Thoms and Bulat Ibragimov
arXiv
arXiv
08/24/2026
DOI: 10.48550/arxiv.2608.23836
url
https://doi.org/10.48550/arxiv.2608.23836View
Preprint (Author's original) This preprint has not been evaluated by subject experts through peer review. Preprints may undergo extensive changes and/or become peer-reviewed journal articles. Open Access

Abstract

Accurate interpretation of volumetric CT requires efficient navigation of 3D image volumes and attention to diagnostically relevant regions. While eye-tracking has been widely studied in 2D medical imaging, its use for expertise assessment in CT settings remains limited. We propose a gaze-informed transformer framework for expertise classification in thoracic CT. Using a DINOv2 backbone, radiologist fixation patterns are integrated into volumetric feature learning through (1) a learnable log-space bias in self-attention and (2) gaze-weighted pooling of patch embeddings. We trained and evaluated our approach on 182 CT reading sessions from five radiologists with varying levels of experience. On a held-out test set, the model achieves an ROC-AUC of 0.91 and F1 score of 0.86, outperforming adapted methods. These findings suggest that incorporating visual search behavior into transformers may support objective, process-based expertise assessment in radiology. Code is available via https://github.com/leiluk1/GazeToSkill.
Computer Science - Artificial Intelligence Computer Science - Computer Vision and Pattern Recognition Computer Science - Learning

Details

Metrics

1 Record Views
Logo image