arXiv AI By Jim O'Connor, Annika Hoag, Sarah Goyette, Gary B. Parker

Temporal-Difference Learning for Dragonchess

Read the original on arXiv AI →

The paper examines the performance of evolutionary transfer learning and TD(lambda) in the three‑dimensional chess game Dragonchess. By re‑implementing the engine in C++ to accelerate play, the authors ran 10,000 games with statistical confidence, showing both adaptive methods outperform all other agents in a round‑robin tournament. The results indicate no significant performance difference between the evolved and learned evaluation functions, demonstrating the effectiveness of adaptive techniques in complex, novel game domains.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 23

PTCG-Bench: Can LLM Agents Master Pok\'emon Trading Card Game?

PTCG-Bench is a new benchmark that uses the Pokémon Trading Card Game to evaluate large language model (LLM) agents on two fronts: their decision‑making within a single complex game environment and their capacity to evolve through accumulated experience. The benchmark includes a modular harness ablation to isolate agent performance from model capability. Experiments show that while LLM agents can achieve non‑trivial gameplay, sustained self‑evolution remains difficult and performance depends on harness design.

By Dongdong Hua, Yifei Sun, Renhong Huang, Feng Gao, Chunping Wang, Yang Yang
arXiv Machine Learning
Sep 4

LLM-Guided Reinforcement Learning for Adaptive NPC Behavior in Multi-Agent Combat Games

The paper explores a runtime strategy-selection framework where a large language model (LLM) guides a pre‑trained reinforcement learning (RL) policy for non‑player characters (NPCs) in a Unity combat game without altering the underlying policy. Five NPC agents sharing a PPO policy were compared in a baseline setup and an LLM‑augmented setup, where a locally hosted Mistral 7B model assigns one of four tactical tags every five seconds based on live game state. Across 600 episodes against three scripted opponents, the LLM‑augmented agents more than doubled their win rate against a Balanced opponent, improved performance against an Evasive opponent, but struggled against an Aggressive opponent due to over‑reliance on encirclement; analysis of 2,430 strategy selections revealed limited zero‑shot differentiation with the model favoring Surround in 83.8% of cases.

By Hrithika Deepu Nair, Kayvan Karim