The study evaluates automatic tooth segmentation on panoramic radiographs using a large annotated corpus of 1,422 images and 42,142 tooth polygons. It finds that increasing input resolution improves boundary precision (mask mAP50‑95 rises from 0.656 to 0.717) while detection performance remains unchanged, and that architectural changes have minimal impact on in‑domain accuracy. Targeted interventions such as LoRA adaptation, promptable foundation models, and anatomical label assignment provide negligible gains, indicating that resolution and acquisition diversity should be prioritized over model novelty.
By Muhammad Rehan, Moaz Amjad, Syed Danial Ahmed, Mariam Adnan, Haider Ali
arXiv:2606. 02914v1 Announce Type: new Abstract: Background: Oral diseases affect nearly 3.
By Sema Helali, Lina Abu Nadab, Sausan Alqawas, Alaa Abd-Alrazaq, Faleh Tamimi, Rafat Damseh
arXiv:2609.14703v1 Announce Type: new
Abstract: Dental caries and endodontic disease are among the most common health conditions worldwide, and intraoral periapical radiographs are central to their d...
By Md Jubaer Rahman, Ulas Bagci
Dental caries and endodontic disease are among the most common health conditions worldwide, and intraoral periapical radiographs are central to their detection, treatment planning, and follow-up. Auto...
Current evaluation protocols for Vision-Language Models (VLMs) in Radiology Report Generation (RRG) rely on report-level metrics that measure lexical overlap or aggregate clinical correctness. However, such metrics do not test whether individual diagnostic statements stem from the actual pathological evidence visible in the image.
arXiv:2608. 03890v1 Announce Type: cross Abstract: A clinically useful chest X-ray system must go beyond fluent report generation: it should classify findings with tunable decision thresholds, localize them spatially, and derive the anatomical measurements upon which many diagnoses depend.
By Mercy Prasanna Ranjit, Anirban Porya, Sathvik Joel, Niharika Vadlamudi, Nikhilesh Chowdary Eathamukkala, Prasanth V V, Abhyuday Kumara Swamy, Pranay Narhari Umredkar, Pradeep Narayan, Vivek Rajagopal, Tanuja Ganu
arXiv:2604.16729v2 Announce Type: replace-cross
Abstract: State-of-the-art large language models (LLMs) show high performance in general visual question answering. However, a fundamental limitation r...
By Ayhan Can Erdur, Daniel Scholz, Jiazhen Pan, Benedikt Wiestler, Daniel Rueckert, Jan C. Peeken
arXiv:2609.32352v2 Announce Type: replace-cross
Abstract: Vision-language models (VLMs) have shown increasing potential for medical image understanding, yet their capabilities in ophthalmic imaging r...
By Gujie Shao, Zixun Xie, Xuechun Xing, Ruixiang Wang, Ziyun Lan, Yanlin Qi, Gangyi Zhang, Yuxin Yang, Dawei Li, Haiming Tang
arXiv:2608.21140v1 Announce Type: cross
Abstract: Reliable spatial understanding is an important prerequisite for future medical vision-language systems that aim to support radiological report genera...
By Simon Vincent Abel, Heiko Hillenhagen, Michael G\"otz, Timo Ropinski, Ayhan Can Erdur, Daniel Santak Wolf
arXiv:2606. 17710v1 Announce Type: cross Abstract: Medical vision-language models report strong chest radiograph accuracy, and this is increasingly read as evidence that they use the image.
By Mahshad Lotfinia, Sebastian Ziegelmayer, Lisa Adams, Daniel Truhn, Andreas Maier, Soroosh Tayebi Arasteh
The paper introduces MedREAL, a unified framework that aligns linguistic reasoning with spatial grounding for medical visual question answering and segmentation. MedREAL employs Seg Anchored Reasoning Pooling (SARP) to extract semantic evidence from segmentation tokens and a Reasoning-to-Visual (R2V) fusion mechanism to integrate these features into a segmentation pipeline. Using the newly created MedRAVS-13K dataset, MedREAL achieves superior performance, reporting 68.49% gIoU and 70.47% cIoU, and generates evidence masks that consistently match textual diagnoses.
By Haowen Gu, Gensheng Pei, Junzhu Mao, Qiong Wang, Mingwu Ren, Yazhou Yao
Prototype-based networks provide inherently interpretable classification by linking predictions to learned exemplars, but their use in 3D point clouds and clinical surface-pair reasoning remains limited. We introduce ProtoPointNet, a prototype-based model for dental occlusion classification from registered upper--lower intraoral arch pairs.