arXiv AI

A doctrine-grounded visual question answering dataset for Tactical Combat Casualty Care

arXiv AI
Sep 21

LiteMedCoT-VL: Parameter-Efficient Adaptation for Medical Visual Question Answering

LiteMedCoT-VL is a parameter‑efficient pipeline that transfers chain‑of‑thought reasoning from a 235B teacher model to a 2B student model using LoRA fine‑tuning on explanation‑enriched data. The approach enables a compact vision‑language model to perform medical visual question answering without relying on image captions, achieving 64.9% accuracy on the PMC‑VQA benchmark—an 11‑point improvement over the zero‑shot Qwen3‑VL‑4B baseline. Visual grounding analysis confirms that the model bases its predictions on image content rather than textual priors.

By Runze Ma, Shunbo Jia, Haonan Lyu, Guo Liu, Caizhi Liao
arXiv AI
Jul 9

Overview of the NLPCC 2026 Shared Task 1: Difficulty-Aware Multilingual and Multimodal Medical Instructional Video Understanding Evaluation

arXiv:2607. 06618v1 Announce Type: cross Abstract: Following the CMIVQA, MMI-VQA, and M4IVQA challenges in NLPCC 2023--2025, we introduce the Difficulty-Aware Medical Instructional Video Question Answering (DA-MIVQA) shared task for NLPCC 2026.

By Shenxi Liu, Kan Li, Mingyang Zhao, Yuhang Tian, Bin Li
arXiv AI
Jul 21

MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering

arXiv:2604. 09757v2 Announce Type: replace-cross Abstract: Medical vision--language models (VLMs) have shown strong potential for medical visual question answering (VQA), yet their reasoning remains largely text-centric: images are encoded once as static context, and subsequent inference is dominated by language.

By Suyang Xi, Songtao Hu, Yuxiang Lai, Wangyun Dan, Yaqi Liu, Shansong Wang, Xiaofeng Yang
Hugging Face Trending Papers
Jun 24

SurgAtlas: A Large-Scale Surgical Video-Language Dataset with 2,391 Hours of Open and Minimally Invasive Surgery

We introduce SurgAtlas, the largest surgical video-language dataset to date, comprising 15,291 videos (2,391 hours) spanning 18 surgical specialties and over 5,000 procedure types, sourced entirely from publicly available YouTube content. SurgAtlas is also the first surgical video-language dataset to include open surgery at scale, with 6,182 open procedure videos alongside over 9,000 minimally invasive recordings, and the first to establish standardized benchmarks for open-surgery video understanding.

arXiv AI
Jun 11

OpenMedReason: Scientific Reasoning Supervision for Medical Vision-Language Models

arXiv:2606. 12169v1 Announce Type: cross Abstract: High-stakes clinical use of large vision-language models (LVLMs) requires reasoning that is grounded in visual evidence and clinical knowledge, not just correct final answers.

By Negin Baghbanzadeh, Pritam Sarkar, Michael Colacci, Abeer Badawi, Adibvafa Fallahpour, Arash Afkanpour, Leonid Sigal, Ali Etemad, Elham Dolatabadi
arXiv AI
Aug 11

Med-CRAFT: An Information System for Explainable and Configurable Construction of Multimodal Medical QA Datasets

arXiv:2512. 01045v2 Announce Type: replace Abstract: Data-intensive artificial intelligence applications increasingly rely on large-scale, high-quality, explainable, and reproducible datasets, yet the construction of such datasets often remains labor-intensive, weakly traceable, and difficult to configure.

By Shenxi Liu, Kan Li, Mingyang Zhao, Yuhang Tian, Bin Li
arXiv Computer Vision
Sep 4

SurgAtlas: A Large-Scale Surgical Video-Language Dataset with 2,391 Hours of Open and Minimally Invasive Surgery

SurgAtlas is the largest surgical video‑language dataset, containing 15,291 videos (2,391 hours) across 18 specialties and over 5,000 procedure types, all sourced from public YouTube. It uniquely includes open‑surgery videos at scale (6,182) alongside more than 9,000 minimally invasive recordings, and introduces standardized benchmarks for open‑surgery video understanding. The dataset offers a rich, multi‑tier annotation schema—segment‑level captions, step/phase descriptions, video‑level surgical narratives, and reasoning‑oriented VQA pairs—validated by experts and built through an automated LLM‑enriched pipeline. "whyItMatters":"SurgAtlas provides an unprecedentedly large, diverse, and clinically validated resource that can train and benchmark multimodal surgical AI models, advancing the development of next‑generation foundation models for surgery."

By Filippos Bellos, Andre S. Gala-Garza, Miaowei Wang, Alyssa M. Hardin, Ahmad M. Hider, Li Yayuan, Jing Bi, Susan Liang, Chenliang Xu, Donald S. Likosky, Jason J. Corso
arXiv Computer Vision
Aug 25

SurgTEMP: Temporal-Aware Surgical Video Question Answering with Text-guided Visual Memory for Laparoscopic Cholecystectomy

arXiv:2603.29962v4 Announce Type: replace Abstract: Surgical procedures are inherently complex and risky, requiring extensive expertise and constant focus to navigate evolving intraoperative scenes....

By Shi Li, Vinkle Srivastav, Nicolas Chanel, Saurav Sharma, Nabani Banik, Lorenzo Arboit, Kun Yuan, Pietro Mascagni, Nicolas Padoy
arXiv Computation and Language
Sep 11

BodyCam-VQA: Enhanced Body-Worn Camera Video Captioning via Multimodal Reasoning and Probe Question Generation

The paper introduces BodyCam-VQA, an adaptive visual question answering framework designed to improve captioning of police body‑worn camera footage. By employing structured multimodal reasoning and probe question generation, the system extracts fine‑grained visual evidence that conventional captioning models miss. Experiments with various question generation models show that this VQA‑driven approach yields more reliable, objective, and detailed records of enforcement events.

By Karish Gupta, Matthew Alex, Alex Li, Yang Wu, Yun-Wei Chu, Kashif Munir, Xiaotian Zhou, Zhengping Ji, Xiaozhong Liu
arXiv AI
Jul 21

Med-OPD: Improving Medical Vision-Language Models via Evidence-Aware On-Policy Distillation

arXiv:2607. 16303v1 Announce Type: cross Abstract: Medical Vision-Language Models (Med-VLMs) require reliable reasoning from fine-grained visual evidence, yet existing models can produce plausible clinical answers by relying on language priors or medical templates rather than truly attending to diagnosis-critical regions.

By Yunhang Qian, Jiaquan Yu, Jiawei Liu, Meng Wang, Hongwei Bran Li, Xiaobin Hu