arXiv:2601. 03309v2 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models, which integrate pretrained large Vision-Language Models (VLM) into their policy backbone, are gaining significant attention for their promising generalization capabilities.
By Jianke Zhang, Xiaoyu Chen, Qiuyue Wang, Mingsheng Li, Yanjiang Guo, Yucheng Hu, Jiajun Zhang, Shuai Bai, Junyang Lin, Jianyu Chen
arXiv:2506. 14126v2 Announce Type: replace-cross Abstract: Modern deep learning is increasingly characterized by the use of open-weight foundation models that can be fine-tuned on specialized datasets.
By Stefan Horoi, Guy Wolf, Eugene Belilovsky, Gintare Karolina Dziugaite
The paper introduces Mask Fine‑Tuning (MFT), a new approach for adapting Vision‑Language Models that avoids modifying backbone weights. MFT learns masks to selectively route information through existing pretrained connections, dynamically uncovering subnetworks that better align with downstream tasks. Experiments demonstrate that MFT consistently outperforms both Full Fine‑Tuning and Parameter‑Efficient Fine‑Tuning across multiple benchmarks, while also offering insights into how pretrained VLMs reorganize their internal pathways during adaptation.
By Mingyuan Zhang, Yue Bai, Yifan Wang, Yiyang Huang, Yun Fu
Knowledge-Intensive Visual Question Answering (KI-VQA) benchmarks evaluate Vision-Language Models (VLMs) as multimodal knowledge assistants by requiring external information beyond a provided image to answer questions. KI-VQA involves multiple sub-problems -referring expression understanding, visual grounding, object recognition, knowledge retrieval, and reasoning-yet existing benchmarks typically report only end-task accuracy, obscuring where failures arise.
arXiv:2410.21582v4 Announce Type: replace-cross
Abstract: Large-scale pretrained models are widely leveraged as foundations for learning new specialized tasks via fine-tuning, with the goal of mainta...
By Jaedong Hwang, Brian Cheung, Zhang-Wei Hong, Akhilan Boopathy, Pulkit Agrawal, Ila Fiete
arXiv:2608. 15580v1 Announce Type: new Abstract: Reliable endoscopic polyp reporting requires integrating quantitative lesion sizing, standardized Paris classification, and clinically meaningful morphological description within a single record.
By Ruijie Yang, Yan Zhu, Peiyao Fu, Siyuan Li, Te Luo, Zhihua Wang, Quanlin Li, Pinghong Zhou, Xian Yang, Shuo Wang