arXiv:2608.29374v1 Announce Type: new
Abstract: Tool-augmented multimodal reasoning integrates external tools (e.g., object detection, depth estimation) into multimodal large language models (MLLMs)...
By Yu Cheng, Arushi Goel, Hakan Bilen
The paper introduces an interventional protocol to assess how vision‑language models (VLMs) explain the impact of missing modalities on their predictions. By comparing the models’ self‑explanations with actual changes observed after restoring missing inputs, the study finds that VLMs routinely overstate the sufficiency of available evidence and underestimate the effect of adding back missing modalities. Across eight open‑weight VLMs and four tasks, the discrepancy between predicted and realized changes is substantial, revealing systematic mischaracterization of modality dependence.
By Aydin Javadov, Daniel Schoess, Florian von Wangenheim
arXiv:2507. 02778v3 Announce Type: replace-cross Abstract: Although large language models (LLMs) have transformed AI, they still make errors and follow unproductive reasoning paths.
By Ken Tsui
Self-improvement for multimodal large language models (MLLMs) is typically driven by reward-based methods that provide only coarse scalar feedback. Distillation offers a richer alternative through dense token-level supervision, but in the visual domain it usually depends on privileged context constructed using external annotations and tools, or stronger models.
arXiv:2607. 23125v1 Announce Type: new Abstract: Post-training enables vision-language models (VLMs) to understand human instructions and perform various downstream tasks.
By Shuai Wang, Daoan Zhang, Zhe Tang, Hao Cheng, Jiaheng Wei
arXiv:2608.22857v1 Announce Type: new
Abstract: Vision-language models (VLMs) frequently fail at visual change reasoning, even when their vision encoders contain sufficient information. We observe th...
By Youdi Li