arXiv AI By Yutong Bian, Dongjie Cheng, Heming Xia, Yongqi Li, Wenjie Li

Optical Reasoning: Rethinking Images as an Expressive Reasoning Medium Beyond Text

Read the original on arXiv AI →

arXiv:2606. 09585v1 Announce Type: new Abstract: Chain-of-Thought (CoT) improves the performance of Large Language Models (LLMs) and has been extended to Multimodal Large Language Models (MLLMs).

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

arXiv AI
Jun 11

MentisOculi: Revealing the Limits of Reasoning with Mental Imagery

arXiv:2602. 02465v2 Announce Type: replace Abstract: Frontier models are transitioning from multimodal large language models (MLLMs) that merely ingest visual information to unified multimodal models (UMMs) capable of native interleaved generation.

By Jana Zeller, Thadd\"aus Wiedemer, Fanfei Li, Thomas Klein, Prasanna Mayilvahanan, Matthias Bethge, Felix Wichmann, Ryan Cotterell, Wieland Brendel