arXiv AI By Linzhe Jiang, Jiayuan Huang, Changhao Zhang, Chunyang Jiang, Zhehua Mao, Mobarak I. Hoque

RefineRank: Joint Box Refinement and Ranking for Surgical Spatio-Temporal Grounding

Read the original on arXiv AI →

RefineRank introduces a lightweight module, RefineNet, that jointly refines bounding boxes and ranks them for surgical spatio‑temporal grounding. By combining frozen medical vision‑language features with proposals from a frozen open‑set detector, it predicts coordinate corrections and quality scores for each candidate box, selecting the best refined or original box. On MedVidBench, RefineRank achieves the highest reported STG mIoU of 0.421 and improves oracle bounds and overall mIoU in controlled evaluations.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 29

Why Does Grounding Hurt Medical VQA? Benchmarking, Diagnosis, and Fine-Tuning of Vision-Language Models

arXiv:2604. 27720v2 Announce Type: replace Abstract: Vision-language models (VLMs) are increasingly applied to medical visual question answering (Med-VQA), yet whether they can \emph{localize} the evidence behind their answers---a prerequisite for clinical auditability---is poorly characterized.

By Xupeng Chen, Binbin Shi, Chenqian Le, Qifu Yin, Lang Lin, Haowei Ni, Ran Gong, Panfeng Li
arXiv AI
Aug 17

MedClaw: Heuristic Agent Harness for Long-Horizon Surgical Video Reasoning

arXiv:2608. 14015v1 Announce Type: cross Abstract: Understanding tens-of-minutes surgical videos requires long-horizon temporal reasoning, answering what happens before, after, or across stages of a procedure by grounding the question in visual evidence spread across time.

By Yingying Fan, Penghui Du, Leyan Zhu, Runze He, Zimeng Wu, Yuxuan Zhang, Liang Chen, Jiahao Xie, Jiangtang Wang, Shuai Shao, Anchao Yang, Yutong Bai, Yan Wang
arXiv Computer Vision
Aug 25

SurgTEMP: Temporal-Aware Surgical Video Question Answering with Text-guided Visual Memory for Laparoscopic Cholecystectomy

arXiv:2603.29962v4 Announce Type: replace Abstract: Surgical procedures are inherently complex and risky, requiring extensive expertise and constant focus to navigate evolving intraoperative scenes....

By Shi Li, Vinkle Srivastav, Nicolas Chanel, Saurav Sharma, Nabani Banik, Lorenzo Arboit, Kun Yuan, Pietro Mascagni, Nicolas Padoy
arXiv Computer Vision
Sep 17

Lumen: Parameter-Efficient Alignment of Pretrained Vision and Language Encoders for Zero-Shot Computational Pathology

Lumen is a pathology vision‑language model that aligns frozen unimodal foundation models (Virchow2 and BioMedBERT) using rank‑4 adapters and projection heads, training only 0.40% of the total parameters on the QUILT‑1M corpus. It achieves the highest mean chance‑corrected balanced accuracy (0.546) across nine zero‑shot patch benchmarks and demonstrates strong performance on lymph‑node metastasis detection, with AUROC scores of 0.964 internally and 0.955 externally. While it ranks third in cross‑modal retrieval, Lumen’s low‑parameter training yields competitive results at both patch and slide levels.

By Kiarash Tajbakhsh, Abdelrahman Faqieh, Michael Jopiti, Javier Garcia-Baroja, Philipp Zens, Branislav Zagrapan, Yuri Tolkach, Martin D. Berger, Aurel Perren, Bastian Dislich, Inti Zlobec, Amjad Khan