arXiv:2601. 03309v2 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models, which integrate pretrained large Vision-Language Models (VLM) into their policy backbone, are gaining significant attention for their promising generalization capabilities.
By Jianke Zhang, Xiaoyu Chen, Qiuyue Wang, Mingsheng Li, Yanjiang Guo, Yucheng Hu, Jiajun Zhang, Shuai Bai, Junyang Lin, Jianyu Chen
arXiv:2609.15131v1 Announce Type: cross
Abstract: Multimodal large language models (MLLMs) require substantial computation to process numerous visual tokens across all transformer layers. Most method...
By Yuyao Sun, Tao Deng, Shuang Li, Deqing Wang
arXiv:2507.00754v3 Announce Type: replace
Abstract: The integration of Large Language Model (LLMs) blocks with Vision Transformers (ViTs) holds immense promise for vision-only tasks by leveraging the...
By Selim Kuzucu, Muhammad Ferjad Naeem, Anna Kukleva, Federico Tombari, Bernt Schiele
arXiv:2607. 23125v1 Announce Type: new Abstract: Post-training enables vision-language models (VLMs) to understand human instructions and perform various downstream tasks.
By Shuai Wang, Daoan Zhang, Zhe Tang, Hao Cheng, Jiaheng Wei
Hidden‑Shot introduces an implicit prompt mechanism that extracts task‑specific visual information and merges it with in‑task processing to boost one‑shot performance on new low‑level vision tasks. The method injects this prompt cost‑effectively while minimally altering the base generalist model’s architecture. A data‑driven evaluation framework, C/U assessment, is proposed to systematically test generalization across conventional and unconventional tasks, and experiments on seven and ten datasets show Hidden‑Shot outperforming state‑of‑the‑art models.
By Shao-Jun Xia, Xianzheng Ma, Zichong Meng
arXiv:2607. 28967v1 Announce Type: cross Abstract: Prompt tuning adapts vision--language models with few trainable parameters, but existing approaches trade off efficiency and adaptation: static textual prompts can overfit source classes, image-conditioned prompts add per-instance computation, and multimodal tuning modifies the visual branch.
By Pouya Parsa, Raoof Zare Moayedi, Seongjin Choi