arXiv AI

MedPruner: Training-Free Hierarchical Token Pruning for Efficient 3D Medical Image Understanding in Vision-Language Models

arXiv:2603. 11625v2 Announce Type: replace-cross Abstract: While specialized Medical Vision-Language Models (VLMs) have achieved remarkable success in interpreting 2D and 3D medical modalities, their deployment for 3D volumetric data remains constrained by significant computational inefficiencies.

arXiv Computer Vision
Aug 28

Multi-Image Visual Token Pruning in Large Visual Language Models

The paper introduces Adaptive Visual Token Pruning (AVTP), a training‑free framework that dynamically selects pruning layers and ratios for large vision‑language models (LVLMs) when processing multiple image sequences. By analyzing visual attention distributions across different LVLM architectures, AVTP adapts token retention to image importance, enabling efficient inference without relying on attention‑based computations incompatible with FlashAttention. Experiments show significant speedups—up to 2× for Qwen3VL‑8B—while preserving or even improving accuracy on multi‑image benchmarks.

By Rongyang Zhang, Chengqiang Lu, Cong Li, Hongchao Gu, Tingjia Shen, Xuyang Zhi, Qimeng Wang, Yan Gao, Yi Wu, Yao Hu, Hao Wang, Enhong Chen
arXiv AI
Aug 11

Resolution Meets Reduction: Efficient Visual Context for 3D Radiology Report Generation

arXiv:2608. 08713v1 Announce Type: cross Abstract: Vision-language models offer a promising path toward automating radiology report generation, but applying them to full 3D CT volumes poses substantial computational challenges.

By Jonathan Suprijadi, Raphael Stock, Moritz Langenberg, David Zimmerer, Kim-Celine Kahl, Stefan Denner, Yannick Kirchhoff, Karol Gotkowski, Maximilian Rokuss, Jeremias Traub, Tassilo Wald, Constantin Ulrich, Klaus Maier-Hein
arXiv Machine Learning
Aug 18

Comprehensive language-image pre-training for 3D medical image understanding

arXiv:2510. 15042v3 Announce Type: replace-cross Abstract: In the 3D medical image domain, vision-language pre-training is used to create vision-language encoders (VLEs) that can support radiologists by retrieving patients with similar abnormalities, predicting likelihoods of abnormality, or, with downstream adaptation, generating radiological reports.

By Tassilo Wald, Ibrahim Ethem Hamamci, Yuan Gao, Sam Bond-Taylor, Harshita Sharma, Maximilian Ilse, Cynthia Lo, Olesya Melnichenko, Anton Schwaighofer, Noel C. F. Codella, Maria Teodora Wetscherek, Klaus H. Maier-Hein, Panagiotis Korfiatis, Valentina Salvatelli, Javier Alvarez-Valle, Fernando P\'erez-Garc\'ia
arXiv AI
Sep 2

SinkPruner: Sink-Free Visual Token Pruning for Multimodal Large Language Models

SinkPruner is a training‑free framework that prunes visual tokens for multimodal large language models by first removing high‑norm redundant tokens with a visual sanitizer and then selectively keeping tokens that align with the text query using a text‑guided pruner. The coarse‑to‑fine design reduces attention sink and dispersion, enabling an 89% token reduction while preserving 96.5% of LLaVA‑1.5’s performance and 91.8% of Qwen2.5‑VL’s performance across twelve image‑language and four video‑language benchmarks. The visual sanitizer also improves existing pruning methods, showing strong transferability.

By Shiyu Li, Zi-Yuan Hu, Shijia Huang, Yanyang Li, Yiwu Zhong, Liwei Wang
arXiv AI
Sep 10

STAR-Pro: Stage-Wise Token Adaptive Reduction with Progressive Refinement for Efficient Large Vision-Language Models

arXiv:2609.05916v1 Announce Type: cross Abstract: Large vision-language models (LVLMs) achieve strong multimodal understanding, but the hundreds to thousands of visual tokens they process impose subs...

By Yichen Guo, Tinghao Wang, Qizhe Zhang, Lingbei Meng, Yuan Zhang, Jiajun Cao, Hao Jiang, Chenwei Wu, Jixian Wu, Sixiang Chen, Tao Luo, Hongyang Cheng, Kai Tang, Chenxi Li, Renyuan Li, Xiande Huang, Wenya Wang, Shanghang Zhang
arXiv Computer Vision
Sep 22

MTMed3D: A Multi-Task Transformer-Based Model for 3D Medical Imaging

MTMed3D is a multi-task Transformer-based model that jointly performs 3D detection, segmentation, and classification in medical imaging. It uses a shared Transformer encoder to produce multi-scale features, with separate CNN decoders for each task. Evaluated on BraTS 2018 and 2019, it achieves strong results, especially in detection, while reducing computational cost and inference time compared to single-task models.

By Fan Li, Arun Iyengar, Lanyu Xu
arXiv Machine Learning
Sep 25

Multimodal Routing and Region Refinement for Language-Guided Medical Image Segmentation

Multimodal Routing and Region Refinement for Language-Guided Medical Image Segmentation (MRSeg) is a parameter‑efficient framework that uses frozen ConvNeXt‑Tiny and PubMedBERT encoders to extract multiscale visual features and clinical text tokens. A joint router predicts a sparse mixture over low‑rank adapter bases, enabling separate adaptation for two visual scales and text while keeping feature‑specific parameters distinct. Region Bridge aggregates dense visual tokens into latent regions using text‑derived queries, refines them via self‑attention and text cross‑attention, and redistributes the refined information back to the feature maps, culminating in a multiscale decoder that combines refined semantic features with shallow image evidence. MRSeg achieves state‑of‑the‑art Dice/mIoU scores on QaTa‑COV19 and MosMedData+ with only 7.11 M trainable parameters and 7.60 GFLOPs.

By Md Maklachur Rahman, Md Hasan Al Banna, Saraf Anjum, Assame Arnob, Tracy Hammond