arXiv AI

Large AI Models in Dental Healthcare: From General-Purpose Systems to Domain-Specific Foundation Models

arXiv:2606. 02914v1 Announce Type: new Abstract: Background: Oral diseases affect nearly 3.

arXiv Computer Vision
Sep 17

AgenTeeth: A Model-Agnostic Framework for Suppressing Hallucination in Frozen Vision-Language Models on Dental X-Rays via Tool Evidence Injection

arXiv:2609.17800v1 Announce Type: new Abstract: Vision-language models (VLMs) remain largely unreliable on panoramic dental radiographs and can rely on learned anatomical priors rather than evidence...

By Ahmed Rafid, Fariya Ahmed, Rumman Adib, Mehedi Ahamed, Ajwad Abrar, Tareque Mohmud Chowdhury
arXiv AI
Sep 2

Benchmarking Vision-Language Models for Automated Pathology Diagnosis and Report Generation

arXiv:2609.00866v1 Announce Type: cross Abstract: The rapid advancement of vision-language models (VLMs) has accelerated progress in computational pathology; however, whole-slide image (WSI)-based pa...

By Yumi Lee, Harim Oh, Hyoryung Kim, Minji Kim, Eunsu Kim, Hyeseong Lee, Junya Fukuoka, Andrey Bychkov, Jijgee Munkhdelger, Rajiv Kumar Kaushal, Ayushi Sahay, Rajni Yadav, Bharathi Prabakaran, Sulen Sarioglu, Serdar Balc{\i}, Ilknur Turkmen, Yuri Tolkach, Christian Harder, Julian Westerdorf, Reinhard Buettner, Audun Ljone Henriksen, Sepp De Raedt, Byung Hyun Lee, Sungjin Lim, Joohoon Lee, Gwanghyun Kim, Se Young Chun, Suryakant Singh, Saarthak Kapse, Prateek Prasanna, Kyung A Kim, Yousun Kang, Sehwan Yoo, Sungman Hong, Shubham Innani, Michael Feldman, Spyridon Bakas, Ujjwal Baid, Prasad Dutande, Suhas Gajare, Bhakti Baheti, Serkan S\"okmen, Ece Tu\u{g}ba Cebeci, Ahmet Hal{\i}c{\i}, Musa Balc{\i}, Kardelen Pe\c{c}enek, Srividhya Sainath, Kyongseok Jang, Messi H. J. Lee, Noorul Wahab, Bodong Du, Jiaming Zhang, Qixiang Zhang, Jang-Hwan Choi, Sangjeong Ahn
arXiv Computer Vision
Aug 26

TLNM: Externally Validated Tooth Detection, Numbering and Segmentation from Smartphone Photographs Using Mask R-CNN

The paper presents TLNM, a Mask R‑CNN based system that detects, numbers, and segments teeth in smartphone photographs. It incorporates a masked gray‑world white‑balancing step and an anatomically constrained detection layer to handle patient‑generated variability. Evaluated on internal and external datasets, the model achieved high AP50, PQ, and F1 scores, demonstrating robust performance across diverse populations and imaging conditions.

By Arash Nedaei, Henna Tiensuu, Elina V\"ayrynen, Saujanya Karki, Jaakko Suutala
arXiv AI
Aug 25

Robust Lightweight Deep Learning Models for Oral Cancer Screening

arXiv:2608.21583v1 Announce Type: new Abstract: Oral cancer is a leading cause of mortality in low-to-middle-income countries, where a shortage of specialists delays diagnosis. While point-of-care sc...

By Siddhant Bharadwaj, Aakash Shedsale, Tejashree Subramanya, Mohd. Azfar, Praveen Birur, Debnath Pal, Shankararama Sharma, Anupama Shetty, Rajesh Sundaresan
arXiv AI
Sep 17

Automated Dental Caries Segmentation in Panoramic Radiographs Using Dual-Stage Deep Learning

The paper introduces a dual‑stage deep learning system for detecting dental caries in panoramic radiographs. It first localizes teeth using Faster R‑CNN, then applies U‑Net for pixel‑wise caries segmentation, converting polygon annotations into high‑resolution binary masks. Trained on 3,000 images with both expert and algorithmic labels, the model achieves an IoU of 0.9013, Dice of 0.9482, Recall of 0.9433, and Precision of 0.9774, outperforming existing methods and reducing false positives.

By Jihun Kim, Kyeonghun Kim, Jong-yeol Lee, Yeongseok Seo, Dohyun Chun
arXiv AI
Aug 28

From Reasoning to Pixels: Grounded Medical Multimodal LLMs for VQA and Segmentation

The paper introduces MedREAL, a unified framework that aligns linguistic reasoning with spatial grounding for medical visual question answering and segmentation. MedREAL employs Seg Anchored Reasoning Pooling (SARP) to extract semantic evidence from segmentation tokens and a Reasoning-to-Visual (R2V) fusion mechanism to integrate these features into a segmentation pipeline. Using the newly created MedRAVS-13K dataset, MedREAL achieves superior performance, reporting 68.49% gIoU and 70.47% cIoU, and generates evidence masks that consistently match textual diagnoses.

By Haowen Gu, Gensheng Pei, Junzhu Mao, Qiong Wang, Mingwu Ren, Yazhou Yao
arXiv AI
Jun 30

IMCBench: A benchmark for multimodal LLMs in Image-grounded Medical Conversations

arXiv:2606. 28556v1 Announce Type: new Abstract: Recent advances in large language models and vision-language models have enabled reasoning over multimodal data, offering opportunities for clinical applications such as decision support and triaging.

By Maria Xenochristou, Ashutosh Joshi, Korosh Vatanparvar, Mohammad Abuzar Hashemi, Prasad Kasu, Deepak Bansal, Anchal Nema, Nivedita Wadhwa, Prashams S Jain, Rebecca Abraham, Will Kimbrough, Dilek Hakkani-Tur, Wilko Schulz-Mahlendorf
arXiv Computer Vision
Sep 7

Explainable Multimodal Deep Learning Integrating Imaging and Clinical Data for Oral Potentially Malignant Disorder Detection

The paper presents M2-OPMDNet, a multimodal deep learning framework that combines white‑light and autofluorescence intraoral images with structured clinical data to detect oral potentially malignant disorders (OPMDs). Using a prospectively collected real‑world dataset, the model achieved an AUC of 0.952, outperforming unimodal approaches and improving detection of visually subtle lesions. Explainability was provided through SHAP analysis, showing that clinical variables significantly contributed to risk estimation alongside imaging features.

By Ruilin You, Yihan Wang, Jiabin Chen, Cherie Wink, Petra Wilder-Smith, Rongguang Liang, Bofan Song