arXiv AI

Matilda: Engine-Agnostic Search with Human Policy Guidance

arXiv:2606. 25176v2 Announce Type: replace Abstract: Chess engines have evolved from search-based systems optimized solely for strength to neural policies capable of modeling human decisions across much of the rating spectrum.

arXiv AI
Sep 7

Abstraction Agent

The paper introduces the Abstraction Agent, a zero‑shot pipeline that employs a large language model to automatically generate continuous strategic features from a natural‑language game description, score private states, and cluster them into abstraction buckets without any game‑specific evaluators or training data. The pipeline consists of four phases—feature discovery with calibration anchors, batched private‑state scoring, correlation‑based feature selection, and k‑means clustering—and achieves significant reductions in lifted‑strategy exploitability in heads‑up no‑limit Texas hold’em and outperforms scalar rank baselines in ROVER Trials. The method also transfers to other games such as four‑card Pot‑Limit Omaha, HUNL preflop and flop, and Riichi Mahjong, demonstrating that it can uncover strategic concepts that align with recognized game theory insights.

By Boning Li, Longbo Huang
arXiv AI
Jun 3

SkillDAG: Self-Evolving Typed Skill Graphs for LLM Skill Selection at Scale

arXiv:2606. 03056v1 Announce Type: new Abstract: As LLM agents adopt large skill libraries, selecting the right subset becomes a structural problem rather than a similarity-matching one: skills depend on, conflict with, specialize, or duplicate one another, a structure invisible to both full enumeration and embedding similarity.

By Tong Bai, Zhenglin Wan, Pengfei Zhou, Xingrui Yu, Wangbo Zhao, Yang You, Ivor W. Tsang
arXiv AI
Sep 23

Synthesizing Reactive Character Behaviors for Continuous Games via Programmatic Policy Search

The paper introduces a method for creating reactive character behaviors in continuous games as compact, human‑readable programs. It searches over a domain‑specific language that uses reactive geometric decisions and higher‑order constructs to discretize continuous behavior space, while eliminating redundant program forms through synthesis antipatterns. The approach, called agentic sketching, combines bottom‑up symbolic enumeration with top‑down guidance from a coding agent, and outperforms either technique alone on a benchmark of 14 continuous games.

By Maxim Gumin, Hsueh-Ti Derek Liu, Victor Zordan, Daniel Ritchie
arXiv AI
Sep 4

AgentRM: Enhancing Agent Generalization with Reward Modeling

AgentRM proposes a generalizable reward model to guide LLM-based agents during test-time search, outperforming direct policy fine-tuning. Three reward modeling strategies—explicit, implicit, and LLM-as-a-judge—are explored, and AgentRM improves base policy performance by an average of 8.8 points across nine tasks, surpassing top general agents by 4.0 points. It also shows strong weak-to-strong generalization and can boost specialized agents, with plans to release code for further research.

By Yu Xia, Jingru Fan, Weize Chen, Siyu Yan, Xin Cong, Zhong Zhang, Yaxi Lu, Yankai Lin, Zhiyuan Liu, Maosong Sun