arXiv AI By Zhichao Yang, Yuanze Hu, Haojie Hao, Longkun Hao, Dongshuo Huang, Hongyu Lin, Gen Li, Lanqing Hong, Yihang Lou, Yan Bai

MIRAGE: Mobile Agents with Implicit Reasoning and Generative World Models

Read the original on arXiv AI →

arXiv:2606. 04627v1 Announce Type: new Abstract: Mobile agents are increasingly expected to operate everyday applications from screenshots and language goals, where reliable control requires reasoning over screen affordances, multi-step navigation, and future state changes.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Aug 27

Latent Chain-of-Thought World Modeling for End-to-End Driving

Latent-CoT-Drive (LCDrive) is a vision‑language‑action model for autonomous driving that replaces natural‑language chain‑of‑thought reasoning with a latent language capturing possible outcomes of driving actions. The model interleaves action‑proposal tokens, aligned with the model’s output actions, and world‑model tokens grounded in a learned latent world model to reason about future outcomes. After a supervised cold‑start using ground‑truth future rollouts, LCDrive is further refined with closed‑loop reinforcement learning, achieving faster inference, higher‑quality trajectories, and greater gains from interactive RL than both non‑reasoning and text‑reasoning baselines on a large‑scale end‑to‑end driving benchmark.

By Shuhan Tan, Kashyap Chitta, Yuxiao Chen, Ran Tian, Yurong You, Yan Wang, Wenjie Luo, Yulong Cao, Philipp Krahenbuhl, Marco Pavone, Boris Ivanovic
arXiv Computation and Language
Sep 21

MIRAGE: Multi-Perspective Creative Language Model Reasoning with Reinforcement Learning Guidance

MIRAGE is a new inference-time framework that enhances large language models by using a Selector to choose effective conceptual perspectives and a Reasoner to solve tasks step-by-step, aggregating multiple perspectives when needed. It is inspired by human cognitive flexibility and is designed to improve performance on complex mathematical, scientific, and logical problems. Experiments on GSM8K, MATH500, MMLU-Pro, and Game-of-24 show that MIRAGE outperforms Chain-of-Thought and diverse prompting ensembles, boosting accuracy with minimal inference overhead.

By Arash Lagzian, Srinivas Anumasa, Dianbo Liu