arXiv:2604. 21827v2 Announce Type: replace Abstract: In accomplishing complex tasks, human cognition typically progresses from abstract to concrete (e.
By Nathanael Jo, Zoe De Simone, Mitchell Gordon, Ashia Wilson
arXiv:2609.38604v1 Announce Type: cross
Abstract: Modern LLM agents increasingly tackle complex tasks through interactive, long-horizon exchanges with users, while existing benchmarks generally assum...
By Zheyuan Zhang, Mengyuan Chao, Ke Xiao, Ziyi Chen, Daoan Zhang, Yan Zhang, Yanfang Ye, Wei Xu
arXiv:2608.29610v1 Announce Type: new
Abstract: The current alignment tuning paradigm for Large Language Models (LLMs) prioritizes surface-level behaviors -- fluency, safety, and tonal consistency. W...
By Chenghao Yang
arXiv:2607. 20485v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated remarkable performance on standard benchmarks, yet it remains largely unexplored whether they truly meet user expectations.
By Miaomiao Li, Yang Wang, Bin Liang, Shudong Liu, Zhiwei Zhang, Kam-Fai Wong
The paper introduces a new training paradigm for text-based world models that prioritizes behavior consistency over traditional state consistency metrics. It proposes the Behavior Consistency Reward (BehR), a step-level metric that evaluates how the likelihood of a logged next action changes between real and predicted states using a frozen Reference Agent. Experiments on WebShop and TextWorld demonstrate that BehR-based training improves long-term alignment, reduces false positives in offline evaluation, and yields modest gains in lookahead planning while maintaining or enhancing single-step prediction quality.
By Youling Huang, Guanqiao Chen, Junchi Yao, Lu Wang, Fangkai Yang, Chao Du, ChenZhuo Zhao, Pu Zhao, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang
IDRBench is a benchmark designed to evaluate the interactive capabilities of deep research agents that use large language models. It introduces controlled opportunities for clarification within a common workflow, comparing autonomous and interactive trajectories by measuring task‑specific report alignment and interaction cost. Experiments on 100 tasks with seven LLMs show that interaction consistently improves alignment, though its effectiveness varies depending on the agents’ questions and feedback integration.
By Yingchaojie Feng, Qiang Huang, Xiaoya Xie, Zhaorui Yang, Jun Yu, Wei Chen, Anthony K. H. Tung