SleepVLM: A Rule-Grounded Vision-Language Model for Auditable Sleep Staging
arXiv:2603. 26738v4 Announce Type: replace-cross Abstract: Sleep staging is essential for sleep assessment and disorder diagnosis.
arXiv:2606. 00087v1 Announce Type: cross Abstract: Effective pre-polysomnography screening for obstructive sleep apnea-hypopnea syndrome (OSAHS) requires combining clinical risk factors with visible craniofacial and neck cues.
arXiv:2603. 26738v4 Announce Type: replace-cross Abstract: Sleep staging is essential for sleep assessment and disorder diagnosis.
arXiv:2603. 26738v3 Announce Type: replace-cross Abstract: While automated sleep staging has achieved expert-level accuracy, its clinical adoption is hindered by a lack of auditable reasoning.
arXiv:2608.23313v1 Announce Type: new Abstract: Vision-language model safety benchmarks typically evaluate only final responses: whether a model refuses, warns, or complies. This outcome-level view c...
arXiv:2606. 17710v1 Announce Type: cross Abstract: Medical vision-language models report strong chest radiograph accuracy, and this is increasingly read as evidence that they use the image.
arXiv:2609.32352v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) have shown increasing potential for medical image understanding, yet their capabilities in ophthalmic imaging r...
The paper presents a two‑stage multimodal framework for chest X‑ray interpretation that incorporates radiologist gaze data from the MIMIC‑Eye dataset. Stage 1 introduces a gaze‑token classifier that fuses image patches, bounding‑box masks, transcription embeddings, and fixation maps, and a curriculum‑scheduled loss that improves accuracy and spatial alignment, yielding a 4.4% AUC gain and 13.3% F1 improvement. Stage 2 translates classifier predictions into region‑specific diagnostic sentences using confidence‑weighted keywords, an expert dictionary, and a prompted large language model, boosting clinical‑term BERTScore and ROUGE over keyword baselines.
arXiv:2606. 15129v1 Announce Type: cross Abstract: Color fundus photography (CFP) is the mainstay for large-scale retinal screening, yet its diagnostic capacity is constrained by the lack of depth-resolved structural information.
MultiViewDx is a physician‑validated multimodal instruction dataset that links medical imaging studies with patient context and normalizes heterogeneous reports into an evidence‑linked workflow (evidence → findings → differential discussion → diagnosis). The dataset covers a wide range of imaging modalities and uses a unified image‑text retriever to ensure that instruction synthesis is grounded in source‑supported evidence. Fine‑tuned models on MultiViewDx achieve the highest average accuracy on four MedVQA benchmarks and receive the strongest overall rating on JAMA Clinical Challenge cases, with ablations confirming the importance of case‑level multi‑view organization and evidence‑linked reasoning.
arXiv:2609.32740v2 Announce Type: replace Abstract: Intraoperative anesthesia requires systems to interpret evolving multimodal evidence, recommend timely management, and revise decisions as patient...
arXiv:2607. 25933v1 Announce Type: cross Abstract: Clinical diagnostic evaluation should not only assess whether models can provide correct diagnoses, but also reflect the realities of clinical practice, including progressive disclosure of multimodal information, dynamic updating of diagnostic hypotheses, and continuous refinement of clinical reasoning.
arXiv:2607. 19063v1 Announce Type: new Abstract: Objective Structured Clinical Examinations (OSCEs) are the gold standard for assessing clinical competence, yet scoring remains vulnerable to examiner subjectivity, fatigue, and cognitive bias.
arXiv:2608.22323v1 Announce Type: new Abstract: The application of Large Language Models (LLMs) to diagnostic decision-making has garnered growing interest. However, existing benchmarks largely focus...