SimSkill is a self‑evolving large‑language‑model agent designed for the SUMO traffic simulator. It continuously detects capability gaps, creates and solves environment‑grounded tasks, verifies solutions via an action–critic loop, and stores experiences in episodic, procedural, and semantic memory. Evaluations on two held‑out benchmarks across three LLM backbones show up to a 25‑percentage‑point improvement in verified success, with procedural and semantic memory contributing complementarily.
By Qi Liu, Qinzheng Wang, Can Li, Yiming Bie, Wanjng Ma
arXiv:2606. 31209v1 Announce Type: new Abstract: Interactive traffic simulation is a vital world model for autonomous driving.
By Lingyu Xiao, Zexin Feng, Xintao Yan
arXiv:2604. 17456v2 Announce Type: replace Abstract: Large language model (LLM) agents have shown strong capabilities in long-horizon reasoning, tool use, and decision-making in digital environments, yet extending them to physically grounded systems remains challenging.
By Siqi Lai, Pan Zhang, Yuping Zhou, Jindong Han, Yansong Ning, Hao Liu
SKILL.state is a new runtime architecture for large language model agents that replaces the traditional append‑only conversational history with an explicit, mutable execution state. At each step the model receives only the immutable skill specification, the current structured state, and the latest observation, discarding intermediate reasoning after validating state updates. Experiments across datasets, models, and environments show that SKILL.state improves task accuracy and significantly reduces cumulative token consumption, proving that explicit execution state is a scalable, architecture‑agnostic abstraction for long‑horizon agent skills.
By Sanket Badhe, Priyanka Tiwari, Jonghyun Chung
arXiv:2601. 21754v3 Announce Type: replace Abstract: While Large Language Models (LLMs) excel in language-based agentic tasks, their applicability to unseen, nonlinguistic environments (e.
By Haoyu Wang, Guozheng Ma, Shugang Cui, Yilun Kong, Haotian Luo, Li Shen, Mengya Gao, Yichao Wu, Xiaogang Wang, Dacheng Tao
arXiv:2607. 16900v1 Announce Type: new Abstract: Training API-calling large language model (LLM) agents demands massive amounts of high-quality trajectories.
By Seanie Lee, Sanjoy Chowdhury, Chao Jiang, Cheng-Yu Hsieh, Ting-Yao Hu, Alexander T Toshev, Oncel Tuzel, Raviteja Vemulapalli