MedVision: Benchmarking Quantitative Medical Image Analysis
arXiv:2511. 18676v2 Announce Type: replace-cross Abstract: Current vision-language models (VLMs) in medicine are primarily designed for categorical question answering (e.
The paper introduces the Semantic Tri-view Pipeline, an interpretable system that automatically screens teledermatology photographs for gradability by analyzing epidermal micro-relief across up to three smartphone views. It uses a lightweight DeepLabV3+ model to segment micro-relief fidelity and aggregates the resulting spatial masks with logistic regression, leveraging viewpoint redundancy to improve robustness. Evaluated on the SCIN dataset, the approach raises the AUC from 0.81 to 0.96 on optically clear cases, offering real‑time, privacy‑by‑design feedback to filter ungradable photo sets before clinician review.
arXiv:2511. 18676v2 Announce Type: replace-cross Abstract: Current vision-language models (VLMs) in medicine are primarily designed for categorical question answering (e.
arXiv:2608.21583v1 Announce Type: new Abstract: Oral cancer is a leading cause of mortality in low-to-middle-income countries, where a shortage of specialists delays diagnosis. While point-of-care sc...
arXiv:2603.26483v2 Announce Type: replace Abstract: Medical edge-AI systems must operate under a difficult tension: delivering reliable diagnostic inference while running on devices with limited batt...
arXiv:2607. 09142v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in online medical consultation, yet existing benchmarks remain poorly aligned with real clinical practice.
arXiv:2512. 21414v2 Announce Type: replace-cross Abstract: Recent tool-use frameworks powered by vision-language models (VLMs) improve image understanding by grounding model predictions with specialized tools.
arXiv:2607. 07673v1 Announce Type: cross Abstract: Medicine is inherently multimodal, requiring clinicians to synthesize information across diverse data streams.
arXiv:2606. 25375v2 Announce Type: replace-cross Abstract: With the rapid adoption of generative AI, synthetic medical images pose growing risks, including diagnostic deception and insurance fraud.
arXiv:2604.12411v2 Announce Type: replace Abstract: Segmentation models based on deep neural networks demonstrate strong generalization for medical image segmentation. However, they often exhibit ove...
arXiv:2506.06104v2 Announce Type: replace-cross Abstract: The rising prevalence of chronic wounds, especially in aging populations, presents a significant healthcare challenge due to prolonged hospit...
arXiv:2512.13742v3 Announce Type: replace-cross Abstract: Medical image classifiers detect gastrointestinal diseases well, but they do not explain their decisions. Large language models can generate...
Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D medical images, and evaluation protocols must align with radiologists' clinical practice and provide an accurate, fine-grained and factualness-driven assessment. In this paper, we introduce ClinFusion, a vision-centric MLLM designed for holistic medical understanding that systematically addresses these limitations.
The paper introduces the Cross‑Modal Triage Network (CMTN), a multimodal deep‑learning model that fuses a Swin Transformer V2 visual encoder with a PubMedBERT text encoder to perform severity‑based triage, pathology detection, and generate visual explanations for chest radiographs. Trained on 34,639 image‑text pairs from MIMIC‑CXR‑JPG, the CMTN achieves high ordinal agreement with reference labels (QWK = 0.9341) and excellent pathology detection (macro‑AUROC = 0.9970) while operating with 34 ms latency. However, a blinded clinical audit revealed low agreement with expert radiologists (QWK = 0.1399) and only modest spatial‑semantic concordance in heatmaps, underscoring the gap between algorithmic performance and clinical judgment.