arXiv:2605.27310v2 Announce Type: replace
Abstract: Cross-view spatial reasoning remains a weak spot for vision-language models (VLMs): they reason in language and discard the fine-grained geometry t...
By Qian Yang, Ankur Sikarwar, Huy Le, Le Zhang, Zhuan Shi, Perouz Taslakian, Aishwarya Agrawal
arXiv:2602. 02465v2 Announce Type: replace Abstract: Frontier models are transitioning from multimodal large language models (MLLMs) that merely ingest visual information to unified multimodal models (UMMs) capable of native interleaved generation.
By Jana Zeller, Thadd\"aus Wiedemer, Fanfei Li, Thomas Klein, Prasanna Mayilvahanan, Matthias Bethge, Felix Wichmann, Ryan Cotterell, Wieland Brendel
arXiv:2606. 09585v1 Announce Type: new Abstract: Chain-of-Thought (CoT) improves the performance of Large Language Models (LLMs) and has been extended to Multimodal Large Language Models (MLLMs).
By Yutong Bian, Dongjie Cheng, Heming Xia, Yongqi Li, Wenjie Li
arXiv:2603. 26779v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have demonstrated impressive reasoning capabilities, yet they struggle with spatial tasks that require mental simulation, such as mental rotation.
By Sergio Y. Hayashi, Nina S. T. Hirata
The paper introduces Mosaic, a multi-image visual harness that lets large language‑vision models (MLLMs) construct visual intermediates using ten composable image operations. It evaluates five re‑representation settings on existing multi‑image benchmarks and a new grounding‑focused benchmark, MosaicBench, finding that visual re‑representation benefits tasks requiring precise visual evidence more than those dominated by high‑level semantics. MosaicAgent‑8B is trained via reinforcement learning to compose these operations without demonstration trajectories, demonstrating diverse problem‑solving patterns.
By Gengyuan Zhang, Xiao Han, Xinyu Xie, Tong Liu, Volker Tresp
The paper introduces SVGLM, a framework that integrates scalable vector graphics (SVG) primitives into vision‑language models to enable image generation within reasoning tasks. By treating SVG both as image descriptions and text instructions, SVGLM offers a compact and interpretable method for connecting text and image reasoning. The authors provide a curated SVG‑based image editing dataset and demonstrate strong SVG generation and image‑aware reasoning performance on a mathematical benchmark.
By Sunli Chen, Ding Zhong, Ziqiao Ma, Jiaxin Liu, Zeyuan Yang, Hao Zhang, Lie Lu, Joyce Chai, Chuang Gan
arXiv:2606. 16783v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) excel at visual reasoning but rely on text-based chain-of-thought (CoT), lacking interpretable visual intermediates.
By Zhiqiang Zhou, Junliang Dai, Xu ling
arXiv:2607. 07117v1 Announce Type: cross Abstract: In text-to-image in-context learning (T2I-ICL), a model has to infer a latent compositional pattern from fewshot demonstrations for generating a query image.
By Stepanida Alekseeva, Jenifer Kalafatovich, Seong-Whan Lee
arXiv:2606. 16122v1 Announce Type: new Abstract: Visual thinking should not only sound right; it should show its evidence.
By Junkai Zhang, Yihe Deng, Kai-Wei Chang, Wei Wang
New experimental AI tool helps people explore the context and origin of images seen online.
The paper introduces a progressive training strategy for embodied vision‑language models aimed at reducing spatio‑temporal hallucinations. It first creates a Chain‑of‑Thought dataset that breaks complex reasoning into detailed spatiotemporal steps, then uses supervised pre‑training on this dataset followed by fine‑tuning with weakly‑labeled data. Experiments show the method improves backbone accuracy and narrows the forward‑backward performance gap from over 70% to 6.53%, indicating stronger dynamic reasoning and fewer temporal biases.
By Xiaoda Yang, Shuai Yang, Can Wang, Jingyang Xue, Menglan Tang, Checheng Yu, Xunzhe Zhou, Sashuai Zhou, Tao Jin, Lixin Yang, Xiangyu Yue, Zhou Zhao