arXiv AI By Jaewoo Lee, Archiki Prasad, Justin Chih-Yao Chen, Zaid Khan, Elias Stengel-Eskin, Mohit Bansal

PRInTS: Reward Modeling for Long-Horizon Information Seeking

Read the original on arXiv AI →

arXiv:2511. 19314v2 Announce Type: replace Abstract: Information-seeking is a core capability for AI agents, requiring them to gather and reason over tool-generated information across long trajectories.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
6d ago

Stepwise Intrinsic Rewards for Reasoning in Large Language Models

The paper introduces Stepwise Marginal Information Gain (MIG), an intrinsic process reward that evaluates how each reasoning step of a large language model (LLM) or vision-language model (VLM) improves the likelihood of the reference answer. MIG rewards only new likelihood maxima, preventing duplicate credit, and is combined with outcome, format, and self‑distillation objectives to guide training. Experiments on eight benchmarks show that this method outperforms outcome‑only reinforcement learning and improves accuracy by up to 4.8 points over binary‑reward training, including a 12.6‑point gain on MathVerse and a 12.9‑point advantage on vision‑language transfer at 7B parameters.

By Xiangwei Wang, Wei Wang, Ken Chen, Nanduni Nimalsiri, Sachith Seneviratne, Saman Halgamuge