arXiv AI

Interactive Task Alignment as a POMDP

arXiv:2607. 16412v1 Announce Type: new Abstract: Current benchmarks for language models primarily evaluate execution on fully specified tasks.

arXiv Machine Learning
Sep 3

Beyond State Consistency: Behavior Consistency in Text-Based World Models

The paper introduces a new training paradigm for text-based world models that prioritizes behavior consistency over traditional state consistency metrics. It proposes the Behavior Consistency Reward (BehR), a step-level metric that evaluates how the likelihood of a logged next action changes between real and predicted states using a frozen Reference Agent. Experiments on WebShop and TextWorld demonstrate that BehR-based training improves long-term alignment, reduces false positives in offline evaluation, and yields modest gains in lookahead planning while maintaining or enhancing single-step prediction quality.

By Youling Huang, Guanqiao Chen, Junchi Yao, Lu Wang, Fangkai Yang, Chao Du, ChenZhuo Zhao, Pu Zhao, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang
arXiv AI
Sep 25

IDRBench: Benchmarking the Interactive Capabilities of Deep Research Agents

IDRBench is a benchmark designed to evaluate the interactive capabilities of deep research agents that use large language models. It introduces controlled opportunities for clarification within a common workflow, comparing autonomous and interactive trajectories by measuring task‑specific report alignment and interaction cost. Experiments on 100 tasks with seven LLMs show that interaction consistently improves alignment, though its effectiveness varies depending on the agents’ questions and feedback integration.

By Yingchaojie Feng, Qiang Huang, Xiaoya Xie, Zhaorui Yang, Jun Yu, Wei Chen, Anthony K. H. Tung
arXiv AI
Sep 3

Thinking effort aligns between humans and reasoning models in abductive reasoning

The study examines how the effort expended by large reasoning models (LRMs) compares to that of humans during abductive reasoning tasks. By analyzing reaction times and reasoning traces, the authors find that LRMs and humans exhibit similar patterns of effort and error types. They also demonstrate that decoding strategies allowing models to explore multiple reasoning paths further align the models’ reasoning costs with human effort.

By Henry Arthur
arXiv AI
Aug 12

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information

arXiv:2608. 10692v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed as mobile assistants, where a key challenge is leveraging personal information scattered across multiple applications (apps) to complete user instructions.

By Junjie Ye, Zhuohui Sheng, Shaofan Liu, Yulun Zhu, Wenjie Fu, Dingwei Zhu, Ming Zhang, Yujiong Shen, Weichao Wang, Xin Zhao, Shihan Dou, Tao Gui, Qi Zhang, Xuanjing Huang, Pluto Zhou
arXiv AI
Aug 26

MetaRAG: Belief-Action Aligned Policy Optimization for Agentic RAG

MetaRAG introduces a belief-action aligned policy optimization framework for agentic retrieval-augmented generation (RAG). It incorporates Verify-first Action Generation and Internal Belief Probing to assess whether the current evidence is sufficient before taking an action, and uses a consistency reward gated by answer correctness to guide training. Experiments on seven public QA benchmarks demonstrate that MetaRAG improves the accuracy-efficiency trade-off over existing RL-based agentic RAG baselines, with benefits that transfer across research settings, optimizers, and model backbones.

By Qiuyi Qi, Tian Liang, Jiamu Wang, Jinjian Zhang, Wei Zhou, Pengcheng Zhu, Linjian Mo, Ming Kong, Jie Liu, Qiang Zhu
arXiv AI
Aug 19

KnowSim: Evaluating Information Calibration in LLM Assistants with User Simulators that Learn

KnowSim introduces an evaluation framework that uses a user simulator with explicit knowledge states to assess how well large language models calibrate information to users. The simulator represents knowledge as a graph of Information Units with prerequisite relationships and updates these states based on learning theory. KnowSim computes Knowledge Gain, Delivery Calibration, and Cognitive Overload metrics, and its rankings align with human judgments, outperforming baseline simulators and revealing model performance differences across user knowledge levels.

By Yoonjoo Lee, Hyoungwook Jin, Tae Soo Kim, Shaoyang Zhang, Philippe Laban, Q. Vera Liao