arXiv AI

Synthesizing Reactive Character Behaviors for Continuous Games via Programmatic Policy Search

The paper introduces a method for creating reactive character behaviors in continuous games as compact, human‑readable programs. It searches over a domain‑specific language that uses reactive geometric decisions and higher‑order constructs to discretize continuous behavior space, while eliminating redundant program forms through synthesis antipatterns. The approach, called agentic sketching, combines bottom‑up symbolic enumeration with top‑down guidance from a coding agent, and outperforms either technique alone on a benchmark of 14 continuous games.

arXiv AI
Jul 2

Coachable agents for interactive gameplay

arXiv:2607. 00642v1 Announce Type: new Abstract: Reinforcement learning has proven to be a valuable tool in the creation of advanced AI and robotic systems, contributing to everything from game playing to robotics to foundation models.

By Roberto Capobianco (Sony AI, Zurich, Switzerland), Harm van Seijen (Sony AI, North America, various locations), Nolan D. Bard (Sony AI, North America, various locations), Neil Burch (Sony AI, North America, various locations), Fatima Davelouis (Sony AI, North America, various locations), Josh Davidson (Sony AI, North America, various locations), Alisa Devlic (Sony AI, Zurich, Switzerland), Yunshu Du (Sony AI, North America, various locations), Ishan Durugkar (Sony AI, North America, various locations), Siddhant Gangapurwala (Sony AI, North America, various locations), Daniel Hernandez (Sony AI, North America, various locations), G. Zacharias Holland (Sony AI, North America, various locations), Sahil Jain (Sony AI, North America, various locations), Kenta Kawamoto (Sony AI, Tokyo, Japan), Raksha Kumaraswamy (Sony AI, North America, various locations), Patrick MacAlpine (Sony AI, North America, various locations), Dustin R. Morrill (Sony AI, North America, various locations), Declan Oller (Sony AI, North America, various locations), Francesco Riccio (Sony AI, Zurich, Switzerland), Akanksha Saran (Sony AI, North America, various locations), Craig Sherstan (Sony AI, Tokyo, Japan), Kaushik Subramanian (Sony AI, Zurich, Switzerland), Thomas J. Walsh (Sony AI, North America, various locations), Samuel Barrett (Sony AI, North America, various locations), Kizza N. Frisbee (Sony AI, North America, various locations), Mady Govil (Sony AI, North America, various locations), Johannes G\"unther (Sony AI, North America, various locations), Varun R. Kompella (Sony AI, North America, various locations), James A. MacGlashan (Sony AI, North America, various locations), Maxwell Svetlik (Sony AI, North America, various locations), Michael D. Thomure (Sony AI, North America, various locations), Jaden B. Travnik (Sony AI, North America, various locations), Kevin Waugh (Sony AI, North America, various locations), Elahe Aghapour (Sony AI, North America, various locations), Florian Fuchs (Sony AI, Zurich, Switzerland), Andreanne Lemay (Sony AI, North America, various locations), Shruti Mishra (Sony AI, Zurich, Switzerland), Takuma Seno (Sony AI, Tokyo, Japan), Peter Stone (Sony AI, North America, various locations), Michael Spranger (Sony AI, Tokyo, Japan), Peter R. Wurman (Sony AI, North America, various locations)
arXiv AI
Sep 25

Coding Agents for Generalized Task and Motion Planning Problems

The paper investigates whether large language model–based coding agents can automatically synthesize programs that solve generalized task and motion planning (TAMP) problems across diverse instances. Using Claude Code and Codex, the authors evaluate 980 generated programs on 100 held‑out environments from KinDER and PDDLStream, achieving mean success rates between 56 % and 95 %—higher than hand‑engineered planners and other baselines—while requiring an order of magnitude less computation per instance. The study demonstrates that coding agents can calibrate physical models, test edge cases, and refine strategies, suggesting they are a strong baseline for generalized TAMP.

By Matteo Merler, Bowen Li, Josh Roy, Yichao Liang, Qianwei Wang, Yixuan Huang, Tom Silver
arXiv AI
6d ago

Up and Down the Abstraction Ladder: Code-Based Skills for Language Agents

The paper introduces CodeHack, a library of code-based skills with natural-language descriptions designed to improve language agents in complex environments like NetHack. By allowing agents to invoke reusable skills instead of selecting individual actions, the study shows that skill-based agents nearly triple game progression and cut inference cost by 86% in zero‑shot settings, while still retaining the option to fall back on primitive actions. In reinforcement learning, skill-based agents learn faster, achieving a 7.2× larger average gain in dungeon level within the same training budget.

By Bart{\l}omiej Cupia{\l}, Jens Tuyls, Maciej Wo{\l}czyk, Davide Paglieri, Martin Klissarov, Benjamin Eysenbach, Piotr Mi{\l}o\'s, Karthik R. Narasimhan
arXiv AI
6d ago

SUN: Agentic Robot Policy Learning with Persistent Task Programs

The paper introduces SUN (Semantically UNified) Programs, typed executables that translate grounded relations into optimal control objectives, satisfaction predicates, and learning rewards. Using the Kuafu harness, a foundation model orchestrates scene preparation, verification, residual reinforcement learning, and data generation, repairing candidate programs and calibrating reward weights. Across nine multi‑stage manipulation tasks, Kuafu achieves an 82.03% success rate, outperforms learned baselines, generates demonstrations 10.57× faster than human teleoperation, and transfers zero‑shot to physical Franka and Kinova robots.

By Weiqi Wang, Zhi Li, Yudong Lei, David Martinez, Xiaofeng Gao, Yuxin Jiang, Chenfanfu Jiang, Yingnian Wu, Demetri Terzopoulos, Ran Gong
arXiv AI
Sep 18

UnifiedPlayers: Enhance Tool-Integrated Reasoning in Agentic Reinforcement Learning

UnifiedPlayers is a cooperative framework that jointly adapts planning, execution, and evaluation for tool-integrated reinforcement learning agents. It consists of a Planning Player that generates tasks, an Execution Player that creates multi-turn trajectories with Python tool calls, and an Evaluation Player that builds executable verifiers, all coordinated by role‑specific rewards under GRPO. The approach outperforms prior baselines on mathematical and general reasoning benchmarks and yields a verifier with high adversarial detection accuracy and more discriminative reward signals.

By Wenjie Liao, Liangjie Zhao, Zehong Cao