The paper investigates how pretraining and midtraining enable reward-based learning by providing necessary information and computation. It analyzes sequential state computation and contextual memory, showing that task‑independent source observations resolve ambiguities in reward adaptation. Experiments on pretrained Qwen2.5 checkpoints across eight worlds demonstrate that correct source and first‑operation supervision significantly improve success rates, and that memory replay and independent confirmation further enhance performance.
By Chiwun Yang, Xiaoyu Li
arXiv:2606. 30068v1 Announce Type: new Abstract: Joint-embedding predictive (JEPA-style) objectives learn representations by predicting future latents.
By Ayan Pendharkar
CAFE (Coupled Agent–Feedback Evolution) is a framework that lets a shared‑parameter model alternate between acting as a search agent and as a critic that provides corrective feedback. By learning when to request feedback and how to use it, CAFE trains the agent to recover from its own failures and shapes rewards both online and offline. Experiments on seven search benchmarks show that CAFE outperforms other RL‑based agents, maintains gains on out‑of‑domain tests, and reduces hallucinations, indicating that co‑evolving feedback is essential for self‑improving search agents.
By Boyang Liu, Senjie Jin, Peixin Wang, Zhangyue Yin, Yibo Wang, Yuhao Zhou, Xinbing Liang, Shizheng Zhu, Yuhui Wang, Jingqi Tong, Zhiheng Xi, Jiazheng Zhang, Clive Bai, Clarenceai, Blaze Chen, Tao Gui, Qi Zhang, Xuanjing Huang
arXiv:2602. 05459v2 Announce Type: replace Abstract: Offline goal-conditioned reinforcement learning (GCRL) is typically benchmarked by the best tuned success rate of each method.
By Jan Malte T\"opperwien, Aditya Mohan, Marius Lindauer
arXiv:2609.27532v1 Announce Type: new
Abstract: Long-horizon agentic tasks require an agent to modify an environment through a sequence of tool calls, with success determined by the final state. The...
By Ming Ma, Yi Zhu, Yiran Zhong, Feida Zhu, Chonghan Liu, Pengkun Jiao, Qichao Wang, Yanhao Jia, Tianming Yang, Steven Hoi
The paper investigates the problem of sharing a single critic across multiple parallel environments in reinforcement learning. It shows that when environments assign different expected returns to the same state, a shared critic must reconcile conflicting value targets, which can distort advantage estimates and misguide policy updates. The authors propose a simple fix—providing the critic with the environment index—demonstrating through bandit models and experiments on CartPole, MuJoCo, BipedalWalker, and 16 Procgen games that this conditional critic stabilizes learning and boosts returns, achieving a 40.8% improvement in aggregate normalized return on unseen levels.
By Zhenya Liu, Yang Meng, Zhuokai Zhao, Xuefeng Liu, Yuxin Chen