arXiv:2609.39134v1 Announce Type: new
Abstract: Visual-token compression improves the efficiency of large vision-language models, but can expose failures that full-token evaluation misses. We study a...
By Shilinlu Yan, Bowen Chen, Yuechen Zhang, Zhenhong Zhou, Li Sun, Sen Su
Text-to-image diffusion models enable data-efficient "mimicry" attacks, wherein adversaries fine-tune the model on a handful of public photos to synthesize convincing forgeries of a target individual....
FeatMark is a watermarking framework that protects images from text‑to‑image diffusion model mimicry attacks by embedding small, scene‑consistent micro‑features instead of pixel‑level perturbations. It constructs domain‑specific feature banks, selects executable features, and injects them via mask‑guided concept editing to create highly localized, natural edits. Experiments on VGGFace2, CelebA‑HQ, and WikiArt show FeatMark remains robust against ten strong watermark removal attacks and several adaptive attacks, with minimal impact on perceptual quality and extending to video mimicry scenarios.
By Haoyang Li, Ruoxi Sun, Qingqing Ye, Benjamin Zi Hao Zhao, Yaxin Xiao, Jason Xue, Haibo Hu
Abundant visual information strengthens vision-language model (VLM) perception, yet massive visual tokens raise inference costs. Existing visual token pruning methods rely on similarity-based guidance, which exploits pairwise text-vision and vision-vision token correlations for compression.
arXiv:2512. 08240v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) rely on hundreds of visual tokens, leading to high computational and memory costs.
By Jusheng Zhang, Xiaoyang Guo, Tongyu Mo, Qinhan Lv, Wenhao Chai, Jian Wang, Keze Wang, Liang Lin
arXiv:2509.23928v3 Announce Type: replace-cross
Abstract: Speculative decoding has proven effective for accelerating inference in Large Language Models (LLMs), yet its extension to Vision-Language Mo...
By Zhinan Xie, Peisong Wang, Shuang Qiu, Jian Cheng
arXiv:2509. 22415v3 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) have achieved strong vision-language performance, yet their token-level visual evidence remains difficult to inspect.
By Jiawei Liang, Jianjie Huang, Ruoyu Chen, Xianghao Jiao, Siyuan Liang, Shiming Liu, Xiaochun Cao
arXiv:2606. 06890v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) frequently rely on language priors, producing confident answers that are weakly grounded in visual evidence.
By Runyu Zhou, Qi Zhang, Qixun Wang, Yisen Wang
arXiv:2609.05916v1 Announce Type: cross
Abstract: Large vision-language models (LVLMs) achieve strong multimodal understanding, but the hundreds to thousands of visual tokens they process impose subs...
By Yichen Guo, Tinghao Wang, Qizhe Zhang, Lingbei Meng, Yuan Zhang, Jiajun Cao, Hao Jiang, Chenwei Wu, Jixian Wu, Sixiang Chen, Tao Luo, Hongyang Cheng, Kai Tang, Chenxi Li, Renyuan Li, Xiande Huang, Wenya Wang, Shanghang Zhang
The paper investigates watermarking for large language model outputs, noting that stronger single-layer watermarks reduce token entropy and weaken subsequent layers. It demonstrates that detectability is limited by entropy and that watermark ensembles monotonically lower entropy and the green‑list ratio. The authors propose using weaker single‑layer watermarks to maintain entropy, showing through theory and experiments that this approach improves both detectability and robustness compared to strong baselines.
By Ruibo Chen, Yihan Wu, Xuehao Cui, Jingqi Zhang, Heng Huang
arXiv:2607. 24017v1 Announce Type: cross Abstract: The empirical success of attention mechanism in Multimodal Large Language Models (MLLMs) often obscures its inherent, subtle flaws.
By Pengkun Jiao, Bin Zhu, Jingjing Chen, Yu-gang Jiang
VIVAS is a new Vision‑Language Model pre‑training framework that addresses the lack of fine‑grained visual perception in existing VLMs. It introduces a unified token space and a dense‑structural‑semantic vision tokenizer that expands the textual vocabulary with visual tokens, enabling vision‑language unified autoregressive supervision over both visual details and linguistic content. Trained on 12.4 T tokens, VIVAS achieves state‑of‑the‑art results on 7 tasks and 39 multimodal benchmarks.
By Zhehan Kan, Yubo Zhu, Xinghua Jiang, Zhixiang Wei, Shifeng Liu, Wei Tong, Sheng Zhong, Qingmin Liao, Wenming Yang, Xin Li, Yinsong Liu, Deqiang Jiang, Xing Sun