arXiv Computer Vision

ZAGNet: Zone-Aware Graph Aggregation Network for Patient-Level Lung Ultrasound Diagnosis

ZAGNet is a Zone‑Aware Graph Neural Network that models temporally tracked lung ultrasound findings as graph nodes linked by anatomical zone adjacency, enabling contextual propagation via a graph transformer and a virtual global node for patient‑level diagnosis. It handles missing zones by operating on a flexible graph structure, and outperforms traditional max/mean pooling on a multicenter dataset of 714 subjects, achieving AUCs of 0.803 for consolidation and 0.893 for pleural effusion. The study demonstrates that graph‑based inter‑zone reasoning improves automated patient‑level LUS assessment.

arXiv AI
Sep 25

Clinical Knowledge Graphs for Chest X-Ray Device Reasoning

The paper introduces an uncertainty‑aware clinical knowledge graph for chest X‑ray device reasoning, capturing device instances, tip estimates, placement assessments, provenance, report events, and temporal links as interconnected evidence. The graph builder processes 30,083 studies from 3,255 patients, producing 914,632 evidence nodes and 884,549 typed relationships, while preserving detailed uncertainty and provenance information for each predicted device. The authors also outline typed data contracts, uncertainty representations, abstention rules, report‑image grounding, and longitudinal query mechanisms, though the current analysis is post‑hoc descriptive and does not yet demonstrate clinical utility.

By Harshil Lodhiya
arXiv Computer Vision
Sep 21

Graph-Augmented Topological Internalization with Dual-Stream Classifiers for Medical Report Generation

The paper introduces GDMRG, a Graph-Augmented Dual-Stream Medical Report Generation framework that incorporates a Topological Knowledge Internalization module using a Graph Convolutional Network to encode disease co-occurrence priors. It employs a dual-stream classifier—one branch generating diagnostic prompts under topological constraints and an auxiliary branch dynamically calibrating decision boundaries for imbalanced samples—alongside a Diagnosis-Guided Spatial Attention mechanism to align visual features with clinical semantics. Experiments on MIMIC-CXR show competitive clinical efficacy and natural language fluency, with strong zero-shot performance on IU X-Ray.

By Moyu Tang, Shangkun Sima, Chupei Tang, Junxiao Kong, Di Wang, Tianchi Lu
arXiv Computer Vision
Aug 27

Hierarchical MoE for Multi-Modal ILD Diagnosis

The paper introduces a hierarchical multimodal mixture-of-experts (MoE) model for interstitial lung disease (ILD) classification. It combines a frozen, pre‑trained imaging expert with structured electronic health records (EHR) through a two‑stage gating system: a modality‑level gate weights imaging and EHR predictions, while a sub‑gating module further decomposes the EHR branch into clinically defined feature groups with learned, group‑specific contributions. The approach preserves stable imaging representations, allows input‑dependent clinical weighting, and enhances interpretability across anatomical regions, imaging–EHR utilization, and EHR feature groups, achieving the highest mean AUC (0.8750 ± 0.0443) under strict patient‑level cross‑validation.

By Alec K. Peltekian, Gorkem Durak, Halil Ertugrul Aktas, Carrie Lynn Richardson, Mary Carns, Kathleen Aren, GR Scott Budinger, Anthony J. Esposito, Alexander Misharin, Alok Nidhi Choudhary, Ankit Agrawal, Ulas Bagci
Hugging Face Trending Papers
Jul 13

Analyzing Image Encoder Choices and Graph Homophily in GCN Frameworks for Breast Ultrasound Classification

Breast ultrasound is widely used for screening, yet automated analysis remains challenging due to speckle noise, acquisition variability, and weak separation of benign and malignant cases in standard ultrasound imaging. Graph convolutional networks (GCNs) have recently emerged as a promising approach by leveraging relationships among similar patient samples.

arXiv AI
Sep 7

Cross-modal triage network: a multimodal deep learning framework for severity-based triage and visual explainability in chest radiographs

The paper introduces the Cross‑Modal Triage Network (CMTN), a multimodal deep‑learning model that fuses a Swin Transformer V2 visual encoder with a PubMedBERT text encoder to perform severity‑based triage, pathology detection, and generate visual explanations for chest radiographs. Trained on 34,639 image‑text pairs from MIMIC‑CXR‑JPG, the CMTN achieves high ordinal agreement with reference labels (QWK = 0.9341) and excellent pathology detection (macro‑AUROC = 0.9970) while operating with 34 ms latency. However, a blinded clinical audit revealed low agreement with expert radiologists (QWK = 0.1399) and only modest spatial‑semantic concordance in heatmaps, underscoring the gap between algorithmic performance and clinical judgment.

By Zinah Ghulam, Richa Mittal, Eranga Ukwatta
arXiv AI
Aug 18

FZ-VLM: A Two Stage Florence-Zephyr Vision Language Model Framework for Pulmonary Nodule Characterization and Clinical Decision Making

arXiv:2608. 15004v1 Announce Type: cross Abstract: Lung cancer remains one of the leading causes of cancer-related mortality worldwide, and Computed Tomography (CT) is a primary imaging tool for screening and followup assessment.

By Pramit Dutta, Jenita Manokaran, Richa Mittal, Ryan Appleby, Eranga Ukwatta