arXiv Machine Learning By Ali Janati, Nikita Kuzmin, Rohit Swamy, Charles Niu

Faynt: Scaling and Optimizing Policies for Competitive Melee

Read the original on arXiv Machine Learning →

Faynt is a family of Transformer policies (10M and 75M parameters) that control all 26 characters in Super Smash Bros. Melee from a single checkpoint. After reinforcement learning, the 10M model wins 98.4% of same‑character games against fourteen specialist and multi‑character releases, and defeats a zero‑delay Slippi‑AI model in all 68 evaluated games. The work details architecture, scaling, hyperparameter transfer, supervised pretraining on 840,000 human replays, post‑training curricula, distillation, and efficient inference, and it releases the weights, benchmark suites, and a platform for automated model tournaments.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Aug 3

DungeonBench: A Benchmark for Rules-Rich Tactical Reasoning in Dungeons & Dragons Combat

arXiv:2607. 29577v1 Announce Type: new Abstract: Games and simulators make valuable benchmarks by turning decisions into measurable outcomes, but many current suites under-test rules-rich tactical reasoning: the ability to choose well when geometry, timing, resources, objectives, and rule interactions all matter at once.

By Ismayil Ismayilov, Atakan Kara, Kaan Oktay
arXiv Machine Learning
Sep 4

LLM-Guided Reinforcement Learning for Adaptive NPC Behavior in Multi-Agent Combat Games

The paper explores a runtime strategy-selection framework where a large language model (LLM) guides a pre‑trained reinforcement learning (RL) policy for non‑player characters (NPCs) in a Unity combat game without altering the underlying policy. Five NPC agents sharing a PPO policy were compared in a baseline setup and an LLM‑augmented setup, where a locally hosted Mistral 7B model assigns one of four tactical tags every five seconds based on live game state. Across 600 episodes against three scripted opponents, the LLM‑augmented agents more than doubled their win rate against a Balanced opponent, improved performance against an Evasive opponent, but struggled against an Aggressive opponent due to over‑reliance on encirclement; analysis of 2,430 strategy selections revealed limited zero‑shot differentiation with the model favoring Surround in 83.8% of cases.

By Hrithika Deepu Nair, Kayvan Karim
arXiv AI
Sep 7

CHAMP: Cross-domain Hybrid Architecture for Matchmaking and Prediction in Online Multi-Player Games

CHAMP is a cross‑domain matchmaking framework for Multiplayer Online Battle Arena games that tackles cold‑start, distribution shift, and data‑scarcity issues by using hybrid player profiles and a Domain‑Aware Win‑rate Network (DAWN). DAWN learns mode‑conditioned representations through a shared network, achieving 67.73% win‑rate prediction accuracy and improving match balance in large‑scale A/B tests. The system reduces imbalanced matches, notably cutting 5‑minute kill crushing rates by up to 20.73% for lower‑tier players.

By Kai Wang, Ge Fan, Chaoyun Zhang, Yuyang Jiang, Yuze Liu
arXiv Machine Learning
Aug 27

ShuttleArena: Interpretable Self-Play in Physics-Based Badminton

ShuttleArena is a physics‑based singles badminton self‑play environment that integrates continuous shuttle flight, player interception, structured shot generation, and post‑shot recovery. The policy employs role‑conditioned outputs, allowing interpretable tactical probes through masked interception choices for receivers and factorized hitter actions over shot azimuth, elevation, speed, and recovery target. Evaluation with frozen checkpoints, controlled tactical probes, recovery ablations, qualitative rollouts, and a human‑data sanity check demonstrates competitive performance and reveals that learned recovery behavior is critically important for success.

By Peize Ding