arXiv Machine Learning

A Large-Scale Vision-Language Dataset Derived from Open Scientific Literature to Advance Biomedical Generalist AI

The paper introduces Biomedica, an open-source dataset sourced from PubMed Central that includes over 6 million scientific articles and 24 million image‑text pairs, along with 27 metadata fields and expert human annotations. To facilitate use, the authors provide scalable streaming and search APIs via a web server. They demonstrate the dataset’s value by training embedding models, chat‑style models, and retrieval‑augmented chat agents, all of which outperform previous open systems in their categories.

arXiv Computation and Language
Sep 22

MedGPT-oss: Training a General-Purpose Vision-Language Model for Biomedicine

arXiv:2603.00842v2 Announce Type: replace Abstract: Biomedical multimodal assistants have the potential to unify radiology, pathology, and clinical-text reasoning, yet a critical deployment gap remai...

By Kai Zhang, Zhengqing Yuan, Cheng Peng, Songlin Zhao, Mengxian Lyu, Ziyi Chen, Yanfang Ye, Wei Liu, Ying Zhang, Kaleb E Smith, Lifang He, Lichao Sun, Yonghui Wu
Hugging Face Trending Papers
Jun 29

DEEPMED Search: An Open-Source Agentic Platform for Medical Deep Research with Introspective Verification

Navigating the deluge of heterogeneous medical data, from academic literature (PubMed) to clinical guidelines (Web) and private knowledge bases, remains a critical bottleneck for evidence-based medicine. While commercial black-box tools lack transparency, standard open-source RAG implementations frequently suffer from reasoning drift when handling complex, long-tail queries.

arXiv AI
Sep 15

MedSAM3: Delving into Segment Anything with Medical Concepts

MedSAM-3 is a text‑promptable medical segmentation model that builds on the Segment Anything Model (SAM) by fine‑tuning it with medical images and semantic concept labels. It enables precise anatomical segmentation through open‑vocabulary text descriptions, moving beyond purely geometric prompts. The accompanying MedSAM-3 Agent incorporates multimodal large language models to perform complex reasoning and iterative refinement, and experiments across X‑ray, MRI, ultrasound, CT, and video modalities show it outperforms existing specialist and foundation models.

By Anglin Liu, Xu R. Cao, Yifan Shen, Yi Lu, Xiang Li, Qianqian Chen, Jintai Chen
arXiv AI
Jun 16

Visual-Seeker: Towards Visual-Native Multimodal Agentic Search via Active Visual Reasoning

arXiv:2606. 15231v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have demonstrated impressive capabilities in many visual tasks, but they often struggle with factual grounding when confronted with complex, open-world scenarios.

By Zhengbo Zhang, Changtao Miao, Jinbo Su, Zhaowen Zhou, Chunxia Zhang, Xukai Wang, Ruiqi Liu, Kaiyuan Zheng, Jiansheng Cai, Bo Zhang, Zhe Li, Shiming Xiang, Ying Yan
arXiv Computer Vision
Sep 15

Designing UNICORN: a Unified Benchmark for Imaging in Computational Pathology, Radiology, and Natural Language

arXiv:2603.02790v2 Announce Type: replace Abstract: Foundation models are changing the way we develop medical artificial intelligence. By learning broadly generalizable features across diverse data m...

By Michelle Stegeman (and on behalf of the UNICORN consortium), Lena Philipp (and on behalf of the UNICORN consortium), Fennie van der Graaf (and on behalf of the UNICORN consortium), Marina D'Amato (and on behalf of the UNICORN consortium), Cl\'ement Grisi (and on behalf of the UNICORN consortium), Luc Builtjes (and on behalf of the UNICORN consortium), Joeran S. Bosma (and on behalf of the UNICORN consortium), Judith Lefkes (and on behalf of the UNICORN consortium), Rianne A. Weber (and on behalf of the UNICORN consortium), James A. Meakin (and on behalf of the UNICORN consortium), Thomas Koopman (and on behalf of the UNICORN consortium), Anne Mickan (and on behalf of the UNICORN consortium), Mathias Prokop (and on behalf of the UNICORN consortium), Ewoud J. Smit (and on behalf of the UNICORN consortium), Fr\'ed\'erique Meeuwsen (and on behalf of the UNICORN consortium), Geert Litjens (and on behalf of the UNICORN consortium), Jeroen van der Laak (and on behalf of the UNICORN consortium), Bram van Ginneken (and on behalf of the UNICORN consortium), Maarten de Rooij (and on behalf of the UNICORN consortium), Henkjan Huisman (and on behalf of the UNICORN consortium), Colin Jacobs (and on behalf of the UNICORN consortium), Francesco Ciompi (and on behalf of the UNICORN consortium), Alessa Hering (and on behalf of the UNICORN consortium)
arXiv Machine Learning
Sep 14

LatentVerse: A Framework for Understanding Shared and Modality-Specific Information in Multimodal Latent Representations

LatentVerse is a new framework that provides a web-based visual analytics platform and a command-line interface for analyzing multimodal latent representations. It unifies diagnostics for representation quality metrics and extends analysis to multimodal settings by decomposing embeddings into shared and modality-specific components. The authors evaluate the tool through simulations, real biomedical data analyses, and a user study, demonstrating its utility for interpretable evaluation of foundation model representations.

By Majd Alafrange, Samuel Friedman, John Kitonyo, Sana Tonekaboni, Mahnaz Maddah
arXiv AI
Sep 17

A visual large language foundational model for medical image recognition using clinician-contributed online resources

The paper introduces ThoughtMed-1M, a large-scale medical visual question answering dataset built from de‑identified images and clinician‑generated commentaries, designed to capture structured clinical reasoning and image‑text alignment. Using this dataset, the authors train FOLTMed, a foundational large language model that achieves state‑of‑the‑art performance on 42 medical VQA benchmarks, with a macro accuracy of 85.4% and improved factuality and similarity metrics over existing models.

By Lingxuan Hou, Yuhua Xie, Yue Hu, Yan Zhuang, Junqi Li, Chengzhi Xia, Binh Phu Nguyen, Abubakar Siddique, Minh Nguyen, Yao Hou, Yanju Bao, Kexin Liu, Ke Chen, Jianjun Sun, Zeqi Li, Trung Nguyen, Jiangli Lin