arXiv Machine Learning By Kaixuan Liu, Guojun Xiong, Weinan Zhang, Shengpu Tang

Autoregressive Diffusion World Models for Off-Policy Evaluation of LLM Agents

Read the original on arXiv Machine Learning →

arXiv:2606. 05558v1 Announce Type: new Abstract: Evaluating large language model (LLM) agents in multi-turn interactive environments is expensive and risky, as it requires online environment interaction.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.

arXiv AI
Jun 2

Improving Diffusion Planners by Self-Supervised Action Gating with Energies

arXiv:2603. 02650v2 Announce Type: replace-cross Abstract: Diffusion planners are a strong approach for offline reinforcement learning, but they can fail when value-guided selection favours trajectories that score well yet are locally inconsistent with the environment dynamics, resulting in brittle execution.

By Yuan Lu, Dongqi Han, Yansen Wang, Dongsheng Li