arXiv AI

Offline-Online Curriculum RL for Multimodal Reasoning

arXiv:2607. 23700v1 Announce Type: new Abstract: Multimodal large language models exhibit capabilities on reasoning tasks, yet often produce flawed intermediate steps while yielding correct final answers.

arXiv AI
2d ago

Selection-Based Structured Reasoning: Toward Efficient Multimodal Search Agents

arXiv:2610.01892v1 Announce Type: cross Abstract: Multimodal agents commonly generate free-form reasoning before each action. For small models, limited model capacity can result in lengthy reasoning...

By Feiyu Gavin Zhu, Xiaoyu Zhu, Jiqi Yang, Rui Yang, Arnab Kumar Mondal, Yancheng Wang, Xinke Deng, Jean Oh, Reid Simmons, Joerg Liebelt, Xiang Kong, Zhongyu Jiang
Hugging Face Trending Papers
Aug 8

StructReward: Efficient Structured Process Rewards for Self-Correcting Multimodal Reasoning

Reinforcement learning with verifiable rewards (RLVR) has emerged as an effective approach for improving multimodal reasoning. However, most existing methods evaluate an entire response using a binary reward based only on final-answer correctness, thereby discarding the supervision available in intermediate reasoning steps.

arXiv Computer Vision
Aug 28

Reason in the Words You Speak: Idiolectal Paraphrasing Off-Policy Traces for Reasoning Distillation in VideoLLMs

The paper introduces Echo-GRPO, a method that rewrites privileged reasoning traces into a model’s own idiolect to align off‑policy supervision with the student policy’s vocabulary. By preserving semantics through Dual‑Reference Decoding, Echo‑GRPO mitigates gradient clipping on critical reasoning tokens and improves reasoning distillation. The approach is instantiated as VideoEcho‑R1 for video reasoning, yielding consistent gains across multiple multimodal LLM backbones and benchmarks, and it can be applied as a plug‑in to both RL and supervised fine‑tuning frameworks.

By Ji Soo Lee, Jinyoung Park, Seohyun Lee, Jongha Kim, Joonmyung Choi, Jinsung Yoon, Hyunwoo J. Kim
arXiv AI
Sep 21

Beyond Final Answers: CRYSTAL Benchmark for Transparent Multimodal Reasoning Evaluation

CRYSTAL is a diagnostic benchmark comprising 6,372 multimodal reasoning instances that assess models through verifiable intermediate steps. It introduces two metrics—Match F1 and Ordered Match F1—to evaluate step-level precision, recall, and order. The benchmark, built via a Delphi-inspired pipeline with four independent MLLMs, reveals systematic failures in current models, such as cherry‑picking and disordered reasoning, and proposes a Causal Process Reward and CPR‑Curriculum to improve reasoning performance.

By Wayner Barrios, SouYoung Jin
arXiv Machine Learning
Sep 14

CanvasAnneal: Curriculum Reinforcement Learning for Diffusion Language Models

CanvasAnneal is a curriculum‑guided reinforcement learning framework designed to improve Diffusion Language Models (DLMs) on complex reasoning and tool‑use tasks. It starts training by injecting reasoning traces from a stronger teacher model into the diffusion canvas, then gradually reduces this guidance so the model learns to generate reasoning independently. Experiments on mathematical reasoning and tool‑use benchmarks show that CanvasAnneal outperforms standard diffusion RL methods such as diffu‑GRPO on tasks like MATH500, Countdown, and Tau2, and accelerates reward improvement, though the gains vary by task.

By Blake Olson, Yuhang Song, Emmett McQuinn, Yuan Shangguan