OpenAI Blog

Thinking with images

arXiv AI
Jun 11

MentisOculi: Revealing the Limits of Reasoning with Mental Imagery

arXiv:2602. 02465v2 Announce Type: replace Abstract: Frontier models are transitioning from multimodal large language models (MLLMs) that merely ingest visual information to unified multimodal models (UMMs) capable of native interleaved generation.

By Jana Zeller, Thadd\"aus Wiedemer, Fanfei Li, Thomas Klein, Prasanna Mayilvahanan, Matthias Bethge, Felix Wichmann, Ryan Cotterell, Wieland Brendel
arXiv AI
3d ago

Rethinking Multi-Image Re-Representation in Multi-Image Understanding

The paper introduces Mosaic, a multi-image visual harness that lets large language‑vision models (MLLMs) construct visual intermediates using ten composable image operations. It evaluates five re‑representation settings on existing multi‑image benchmarks and a new grounding‑focused benchmark, MosaicBench, finding that visual re‑representation benefits tasks requiring precise visual evidence more than those dominated by high‑level semantics. MosaicAgent‑8B is trained via reinforcement learning to compose these operations without demonstration trajectories, demonstrating diverse problem‑solving patterns.

By Gengyuan Zhang, Xiao Han, Xinyu Xie, Tong Liu, Volker Tresp
arXiv Computation and Language
Sep 25

Multimodal Thinking with Renderable Programs

The paper introduces SVGLM, a framework that integrates scalable vector graphics (SVG) primitives into vision‑language models to enable image generation within reasoning tasks. By treating SVG both as image descriptions and text instructions, SVGLM offers a compact and interpretable method for connecting text and image reasoning. The authors provide a curated SVG‑based image editing dataset and demonstrate strong SVG generation and image‑aware reasoning performance on a mathematical benchmark.

By Sunli Chen, Ding Zhong, Ziqiao Ma, Jiaxin Liu, Zeyuan Yang, Hao Zhang, Lie Lu, Joyce Chai, Chuang Gan
arXiv AI
Sep 10

A Progressive Training Strategy for Embodied Vision-Language Models to Mitigate Spatio-Temporal Hallucinations

The paper introduces a progressive training strategy for embodied vision‑language models aimed at reducing spatio‑temporal hallucinations. It first creates a Chain‑of‑Thought dataset that breaks complex reasoning into detailed spatiotemporal steps, then uses supervised pre‑training on this dataset followed by fine‑tuning with weakly‑labeled data. Experiments show the method improves backbone accuracy and narrows the forward‑backward performance gap from over 70% to 6.53%, indicating stronger dynamic reasoning and fewer temporal biases.

By Xiaoda Yang, Shuai Yang, Can Wang, Jingyang Xue, Menglan Tang, Checheng Yu, Xunzhe Zhou, Sashuai Zhou, Tao Jin, Lixin Yang, Xiangyu Yue, Zhou Zhao