arXiv Computer Vision By Boya Wang, Ruizhe Li, Chao Chen, Xin Chen

DINO-Med: A Unified Patch-Based Adaptation Framework for Multi-Modal Medical Image Analysis Applied to Liver Fibrosis Staging

Read the original on arXiv Computer Vision →

DINO-Med introduces a patch‑based framework that adapts natural‑image foundation models, specifically DINOv3, to multi‑modal medical imaging. The method uses training‑free registration, automated localization, and mask‑filtered patch extraction to aggregate patch‑level features into subject‑level diagnostics. In liver fibrosis staging, DINOv3 outperforms handcrafted radiomics, ResNet, and SAM‑Med2D features, achieving 78.4% accuracy for mild fibrosis (S1) and 75.8% for cirrhosis (S4) on the CARE 2025 cohort.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

Hugging Face Trending Papers
Sep 10

DINO-Med: A Unified Patch-Based Adaptation Framework for Multi-Modal Medical Image Analysis Applied to Liver Fibrosis Staging

The paper introduces DINO-Med, a patch‑based framework that adapts natural‑image foundation models to multi‑modal medical imaging, specifically for liver fibrosis staging. It processes raw multimodal scans through training‑free registration, automated localization, and mask‑filtered patch extraction, then aggregates patch‑level insights into subject‑level diagnostics. Using the CARE 2025 Liver Track 4 cohort, the DINOv3‑based approach achieved the highest classification accuracy (78.4% for mild fibrosis and 75.8% for cirrhosis) compared to other feature representations.

arXiv Computer Vision
Aug 25

Expert-level vision-language foundation model for real-world radiology and comprehensive evaluation

arXiv:2409.16183v2 Announce Type: replace Abstract: Radiology is a vital and complex component of modern clinical workflow and covers many tasks. Recently, vision-language (VL) foundation models in m...

By Xiaohong Liu, Guoxing Yang, Yulin Luo, Jiaji Mao, Xiang Zhang, Haibo Wang, Zhiyang He, Ming Gao, Shanghang Zhang, Jun Shen, Guangyu Wang
arXiv AI
Aug 17

CMCNet: Aligning Ultrasound Image Embeddings with Textual TI-RADS Representations for Fine-Grained Thyroid Classification

arXiv:2608. 13939v1 Announce Type: cross Abstract: Ultrasound is the primary imaging modality for assessing thyroid nodules, and the ACR TI-RADS framework standardizes diagnosis through five ultrasound feature categories that are aggregated into five risk levels (TR1-TR5).

By Bingxin Yu, Xueli Wang, Jerry Zhou, Wenyan Wang, Li Wen, Lan Huang, Xin Feng, Fengfeng Zhou, Kewei Li
arXiv AI
Sep 10

Semantic Context-aware mOdality fUsion Transformer (SCOUT): A Context-Aware Multimodal Transformer for Concept-Grounded Pathology Report Generation

SCOUT is a concept‑grounded multimodal transformer that generates whole‑slide pathology reports by integrating local histological patterns, whole‑slide context, and expert‑curated diagnostic concepts. It uses evolving visual representations and recursively updated slide‑ and concept‑conditioned representations, with separate attention pathways during decoding that are fused adaptively for each token. Evaluated on TCGA‑BRCA, HistAI, and REG‑2025, SCOUT outperformed existing methods, improving BLEU, METEOR, and ROUGE‑L scores and raising the Clinical Report Quality Score on REG‑2025.

By Suryakant Singh, Saarthak Kapse, Joel Saltz, Prateek Prasanna