arXiv AI By Leila Khaertdinova, Anna Anikina, Claudia Mello-Thoms, Bulat Ibragimov

Predicting Radiologist Expertise from 3D Gaze Patterns During CT Interpretation

Read the original on arXiv AI →

The study introduces a gaze-informed transformer framework that classifies radiologist expertise during thoracic CT interpretation by integrating eye‑tracking data into volumetric feature learning. Using a DINOv2 backbone, the model incorporates a learnable log‑space bias in self‑attention and gaze‑weighted pooling of patch embeddings. Trained on 182 CT reading sessions from five radiologists, it achieved an ROC‑AUC of 0.91 and an F1 score of 0.86, outperforming adapted baseline methods.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Aug 20

Eyes on the Image: Gaze Supervised Multimodal Learning for Chest X-ray Diagnosis and Report Generation

The paper presents a two‑stage multimodal framework for chest X‑ray interpretation that incorporates radiologist gaze data from the MIMIC‑Eye dataset. Stage 1 introduces a gaze‑token classifier that fuses image patches, bounding‑box masks, transcription embeddings, and fixation maps, and a curriculum‑scheduled loss that improves accuracy and spatial alignment, yielding a 4.4% AUC gain and 13.3% F1 improvement. Stage 2 translates classifier predictions into region‑specific diagnostic sentences using confidence‑weighted keywords, an expert dictionary, and a prompted large language model, boosting clinical‑term BERTScore and ROUGE over keyword baselines.

By Tanjim Islam Riju, Shuchismita Anwar, Saman Sarker Joy, Farig Sadeque, Swakkhar Shatabda
arXiv Computer Vision
Aug 25

Expert-level vision-language foundation model for real-world radiology and comprehensive evaluation

arXiv:2409.16183v2 Announce Type: replace Abstract: Radiology is a vital and complex component of modern clinical workflow and covers many tasks. Recently, vision-language (VL) foundation models in m...

By Xiaohong Liu, Guoxing Yang, Yulin Luo, Jiaji Mao, Xiang Zhang, Haibo Wang, Zhiyang He, Ming Gao, Shanghang Zhang, Jun Shen, Guangyu Wang
arXiv AI
Aug 18

Thinking with Gaze: Sequential Eye-Tracking as Visual Reasoning Supervision for Medical VLMs

arXiv:2603. 06697v2 Announce Type: replace-cross Abstract: Vision--language models (VLMs) process images as visual tokens, yet their intermediate reasoning is often carried out in text, which can be suboptimal for visually grounded radiology tasks.

By Yiwei Li, Yifan Zhou, Huaqin Zhao, Zihao Wu, Zhengliang Liu, Xiang Li, Quanzheng Li, Tianming Liu, Lin Zhao