arXiv:2609.17800v1 Announce Type: new
Abstract: Vision-language models (VLMs) remain largely unreliable on panoramic dental radiographs and can rely on learned anatomical priors rather than evidence...
By Ahmed Rafid, Fariya Ahmed, Rumman Adib, Mehedi Ahamed, Ajwad Abrar, Tareque Mohmud Chowdhury
arXiv:2609.00866v1 Announce Type: cross
Abstract: The rapid advancement of vision-language models (VLMs) has accelerated progress in computational pathology; however, whole-slide image (WSI)-based pa...
By Yumi Lee, Harim Oh, Hyoryung Kim, Minji Kim, Eunsu Kim, Hyeseong Lee, Junya Fukuoka, Andrey Bychkov, Jijgee Munkhdelger, Rajiv Kumar Kaushal, Ayushi Sahay, Rajni Yadav, Bharathi Prabakaran, Sulen Sarioglu, Serdar Balc{\i}, Ilknur Turkmen, Yuri Tolkach, Christian Harder, Julian Westerdorf, Reinhard Buettner, Audun Ljone Henriksen, Sepp De Raedt, Byung Hyun Lee, Sungjin Lim, Joohoon Lee, Gwanghyun Kim, Se Young Chun, Suryakant Singh, Saarthak Kapse, Prateek Prasanna, Kyung A Kim, Yousun Kang, Sehwan Yoo, Sungman Hong, Shubham Innani, Michael Feldman, Spyridon Bakas, Ujjwal Baid, Prasad Dutande, Suhas Gajare, Bhakti Baheti, Serkan S\"okmen, Ece Tu\u{g}ba Cebeci, Ahmet Hal{\i}c{\i}, Musa Balc{\i}, Kardelen Pe\c{c}enek, Srividhya Sainath, Kyongseok Jang, Messi H. J. Lee, Noorul Wahab, Bodong Du, Jiaming Zhang, Qixiang Zhang, Jang-Hwan Choi, Sangjeong Ahn
arXiv:2605. 23995v4 Announce Type: replace-cross Abstract: Self-supervised learning (SSL) is increasingly used in medical image analysis to reduce dependence on costly expert annotations by learning transferable representations from unlabeled data.
By Chathura Wimalasiri, Kishor Nandakishor, Marimuthu Palaniswami
The rapid advancement of vision-language models (VLMs) has accelerated progress in computational pathology; however, whole-slide image (WSI)-based pathology report generation remains limited by the sc...
The paper presents TLNM, a Mask R‑CNN based system that detects, numbers, and segments teeth in smartphone photographs. It incorporates a masked gray‑world white‑balancing step and an anatomically constrained detection layer to handle patient‑generated variability. Evaluated on internal and external datasets, the model achieved high AP50, PQ, and F1 scores, demonstrating robust performance across diverse populations and imaging conditions.
By Arash Nedaei, Henna Tiensuu, Elina V\"ayrynen, Saujanya Karki, Jaakko Suutala
arXiv:2608.21583v1 Announce Type: new
Abstract: Oral cancer is a leading cause of mortality in low-to-middle-income countries, where a shortage of specialists delays diagnosis. While point-of-care sc...
By Siddhant Bharadwaj, Aakash Shedsale, Tejashree Subramanya, Mohd. Azfar, Praveen Birur, Debnath Pal, Shankararama Sharma, Anupama Shetty, Rajesh Sundaresan
arXiv:2607. 15380v1 Announce Type: cross Abstract: Electronic health records combine free-text clinical narratives with structured measurements such as vital signs, laboratory values, and comorbidities.
By Ajay Madhavan Ravichandran, Bilgin Osmandoja, Klemens Budde, Klaus Netter, Tobias Strapatsas, Aljoscha Burchardt, Sebastian M\"oller, Roland Roller
The paper introduces a dual‑stage deep learning system for detecting dental caries in panoramic radiographs. It first localizes teeth using Faster R‑CNN, then applies U‑Net for pixel‑wise caries segmentation, converting polygon annotations into high‑resolution binary masks. Trained on 3,000 images with both expert and algorithmic labels, the model achieves an IoU of 0.9013, Dice of 0.9482, Recall of 0.9433, and Precision of 0.9774, outperforming existing methods and reducing false positives.
By Jihun Kim, Kyeonghun Kim, Jong-yeol Lee, Yeongseok Seo, Dohyun Chun
The paper introduces MedREAL, a unified framework that aligns linguistic reasoning with spatial grounding for medical visual question answering and segmentation. MedREAL employs Seg Anchored Reasoning Pooling (SARP) to extract semantic evidence from segmentation tokens and a Reasoning-to-Visual (R2V) fusion mechanism to integrate these features into a segmentation pipeline. Using the newly created MedRAVS-13K dataset, MedREAL achieves superior performance, reporting 68.49% gIoU and 70.47% cIoU, and generates evidence masks that consistently match textual diagnoses.
By Haowen Gu, Gensheng Pei, Junzhu Mao, Qiong Wang, Mingwu Ren, Yazhou Yao
arXiv:2605. 18419v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) can couple visual perception with open-ended clinical reasoning, making them attractive for computational histopathology.
By Franciskus Xaverius Erick, Johanna Paula M\"uller, Bernhard Kainz
arXiv:2606. 28556v1 Announce Type: new Abstract: Recent advances in large language models and vision-language models have enabled reasoning over multimodal data, offering opportunities for clinical applications such as decision support and triaging.
By Maria Xenochristou, Ashutosh Joshi, Korosh Vatanparvar, Mohammad Abuzar Hashemi, Prasad Kasu, Deepak Bansal, Anchal Nema, Nivedita Wadhwa, Prashams S Jain, Rebecca Abraham, Will Kimbrough, Dilek Hakkani-Tur, Wilko Schulz-Mahlendorf
The paper presents M2-OPMDNet, a multimodal deep learning framework that combines white‑light and autofluorescence intraoral images with structured clinical data to detect oral potentially malignant disorders (OPMDs). Using a prospectively collected real‑world dataset, the model achieved an AUC of 0.952, outperforming unimodal approaches and improving detection of visually subtle lesions. Explainability was provided through SHAP analysis, showing that clinical variables significantly contributed to risk estimation alongside imaging features.
By Ruilin You, Yihan Wang, Jiabin Chen, Cherie Wink, Petra Wilder-Smith, Rongguang Liang, Bofan Song