PersonaForge is a user‑simulation framework that generates realistic multi‑turn interactions between users and agentic systems, addressing the gap that most training data assumes single‑turn queries. It uses a four‑dimensional persona space, SOUL‑driven behavioral control calibrated to real‑user statistics, and Reverse Deep Construction from authentic seed queries to create a 6.3K‑record training set and a 138‑task benchmark called PersonaForge‑Bench across 20 professional domains. Experiments with Qwen3.5‑27B show that training with PersonaForge improves composite scores by 4.1%, especially in Task Completion (+6.0%) and Response Quality (+6.8%), while also reducing turns and tool calls, indicating more efficient interactions.
By Hanglong Lv, Dawei Zhu, Lei Li, Bowen Ye, Huaqiu Liu, Yifan Song, Bofei Gao, Weimin Xiong, Jinhao Dong, Chenhong He, Lingpeng Kong, Qi Liu, Tong Yang, Fuli Luo
Agentick is a unified benchmark for sequential decision‑making agents that evaluates RL, LLM, VLM, hybrid, and human agents on 37 procedurally generated tasks across six capability categories, four difficulty levels, and five observation modalities via a single Gymnasium‑compatible interface. It includes a Coding API, oracle reference policies, pre‑built SFT datasets, a composable agent harness, and a live leaderboard. An evaluation of 27 configurations and over 90,000 episodes shows no single approach dominates, with GPT‑5 mini leading overall, PPO excelling in planning and multi‑agent tasks, and the reasoning harness boosting LLM performance by 3‑10×, while ASCII observations outperform natural language.
By Roger Creus Castanyer, Pablo Samuel Castro, Glen Berseth
UniToolCall introduces a unified framework for tool-use in large language model agents, standardizing toolset construction, dataset generation, and evaluation. The framework aggregates over 22,000 tools and creates a hybrid training corpus of more than 390,000 instances by combining ten public datasets with synthetically generated, structurally controlled trajectories. It models diverse interaction patterns—single‑hop vs. multi‑hop, single‑turn vs. multi‑turn, serial vs. parallel execution—and adds an Anchor Linkage mechanism to enforce cross‑turn dependencies, while converting seven public benchmarks into a common Query–Action–Observation–Answer format for fine‑grained evaluation.
By Yijuan Liang, Xinghao Chen, Yifan Ge, Ziyi Wu, Hao Wu, Changyu Zeng, Wei Xing, Xiaoyu Shen
EDGEGEN is a synthetic task generation framework that extracts compliance rules from a tool‑calling agent’s specification to create database‑grounded edge‑case tasks that violate those rules. By combining EdgeGen with existing synthetic data generation methods, it forms a fully automated closed‑loop system that requires no human annotation. Experiments show that finetuning on EdgeGen data improves performance by 2–42 % on the tau2bench airline domain, while harness optimization yields 10–30 % gains over human‑curated and base harnesses for the Gemma‑4‑e4b model.
By Harshavardhan Abichandani, Penny Chong, Jiyuan Shen, Gunraj Singh, Ashutosh Hathidara, Marcus Duigan Xing Yu, Jane Lo, Atin Ghosh, Yipeng Li, Daniel Dahlmeier
arXiv:2608. 06329v1 Announce Type: cross Abstract: Task-oriented conversational agents are evaluated using curated or automatically generated benchmarks, yet benchmark quality is rarely assessed.
By Noam Koren, Roy Bar-Haim, Abigail Goldsteen
arXiv:2606. 11070v1 Announce Type: cross Abstract: Recent advances in reasoning and tool-calling capabilities of large language models (LLMs) have enabled increasingly capable agentic systems.
By Genta Indra Winata, Amartya Chakraborty, Yuzhen Lin, Swasthi P Rao, Shikhhar Siingh, Houhan Lu, Nadia Bathaee, Sriharsha Hatwar, Paresh Dashore, Anmol Jain, Kshitij Tayal, Xiuzhu Lin, Anirban Das, Sambit Sahu, Shi-Xiong Zhang
KC-Bench is a dynamic interactive benchmark designed to evaluate how large language model agents reconcile user instructions, internal knowledge, and real‑time environmental observations. It consists of 238 manually curated multi‑turn tasks that test world‑knowledge conflicts, input inconsistencies, and multi‑source temporal conflicts, using a user simulator, stateful tools, deterministic environment assertions, an open‑source natural‑language evaluator, and human trajectory verification. Evaluation of nine models—including DeepSeek‑V4‑Flash, GLM‑5.2, and MiniMax‑M3—reveals significant cross‑domain variation, with no model reliably handling factual correction, identity consistency, and temporal conflict resolution across all settings, and shows that missed conflicts can propagate to tool calls or synthetic protected‑data flows.
arXiv:2507. 05257v4 Announce Type: replace-cross Abstract: Recent benchmarks for Large Language Model (LLM) agents primarily focus on evaluating reasoning, planning, and execution capabilities, while another critical component-memory, encompassing how agents memorize, update, and retrieve long-term information-is under-evaluated due to the lack of benchmarks.
By Yuanzhe Hu, Yu Wang, Julian McAuley
SKILL.state is a new runtime architecture for large language model agents that replaces the traditional append‑only conversational history with an explicit, mutable execution state. At each step the model receives only the immutable skill specification, the current structured state, and the latest observation, discarding intermediate reasoning after validating state updates. Experiments across datasets, models, and environments show that SKILL.state improves task accuracy and significantly reduces cumulative token consumption, proving that explicit execution state is a scalable, architecture‑agnostic abstraction for long‑horizon agent skills.
By Sanket Badhe, Priyanka Tiwari, Jonghyun Chung
arXiv:2607. 12085v1 Announce Type: new Abstract: Evaluating retail conversational agents requires methods beyond lexical-overlap metrics to assess intent alignment, factuality, helpfulness, clarity, tone, and overall response quality.
By Niranjan Kumar M, Balaji Nagarajan, Karthik Nair, Faysal Satter, Nithin Surendran
arXiv:2606. 15225v1 Announce Type: cross Abstract: Large-scale learner-task interaction data are crucial for intelligent educational systems but are costly to collect and constrained by privacy and learner engagement.
By Weibo Gao, Qi Liu, Linan Yue, Zheng Zhang, Yichao Du, Fangzhou Yao, Ao Yu, Zhenya Huang, Shijin Wang
arXiv:2604. 18543v4 Announce Type: replace Abstract: Constructing environments for training and evaluating claw-like agents remains a manual, human-intensive process that does not scale.
By Xirui Li, Ming Li, Ion Stoica, Cho-Jui Hsieh, Tianyi Zhou