arXiv Machine Learning

ShuttleArena: Interpretable Self-Play in Physics-Based Badminton

ShuttleArena is a physics‑based singles badminton self‑play environment that integrates continuous shuttle flight, player interception, structured shot generation, and post‑shot recovery. The policy employs role‑conditioned outputs, allowing interpretable tactical probes through masked interception choices for receivers and factorized hitter actions over shot azimuth, elevation, speed, and recovery target. Evaluation with frozen checkpoints, controlled tactical probes, recovery ablations, qualitative rollouts, and a human‑data sanity check demonstrates competitive performance and reveals that learned recovery behavior is critically important for success.

arXiv Machine Learning
1d ago

Faynt: Scaling and Optimizing Policies for Competitive Melee

Faynt is a family of Transformer policies (10M and 75M parameters) that control all 26 characters in Super Smash Bros. Melee from a single checkpoint. After reinforcement learning, the 10M model wins 98.4% of same‑character games against fourteen specialist and multi‑character releases, and defeats a zero‑delay Slippi‑AI model in all 68 evaluated games. The work details architecture, scaling, hyperparameter transfer, supervised pretraining on 840,000 human replays, post‑training curricula, distillation, and efficient inference, and it releases the weights, benchmark suites, and a platform for automated model tournaments.

By Ali Janati, Nikita Kuzmin, Rohit Swamy, Charles Niu
arXiv AI
Jun 10

RoboNaldo: Accurate, Stable and Powerful Humanoid Soccer Shooting via Motion-Guided Curriculum Reinforcement Learning

arXiv:2606. 11092v1 Announce Type: cross Abstract: Elite humanoid soccer shooting requires whole-body stability, high-impulse whole-body interactions, and accuracy to targets.

By Yichao Zhong, Yidan Lu, Yuhang Lu, Tianyang Tang, Haoguang Mai, Yixuan Pan, Tianyu Li, Li Chen, Jingbo Wang, Zhongyu Li, Peng Lu, Hongyang Li
arXiv AI
Sep 15

Learning Human-Like Badminton Skills for Humanoid Robots

arXiv:2602.08370v2 Announce Type: replace-cross Abstract: Realizing versatile and human-like performance in high-demand sports like badminton remains a formidable challenge for humanoid robotics. Unl...

By Yeke Chen, Shihao Dong, Xiaoyu Ji, Jingkai Sun, Zeren Luo, Liu Zhao, Jiahui Zhang, Wanyue Li, Ji Ma, Bowen Xu, Yimin Han, Xuanyi Li, Yudong Zhao, Liyun Li, Peng Lu
arXiv AI
Sep 7

Improving Weak World Models Behind Strong Agents in Atari Pong

The paper investigates the discrepancy between strong agents and their underlying weak world models in Atari Pong by reproducing five visual world-model agents and evaluating their frozen models. Closed‑loop rollouts reveal visual and dynamical failures such as ball disappearance and incorrect motion, while zero‑shot model‑based RL policies trained entirely within the frozen models perform poorly compared to the original agents. To address these issues, the authors introduce Concept‑Guided Spatial Regularization (CGSReg), an auxiliary loss focused on task‑critical ball regions, which improves both pixel‑space zero‑shot MBRL performance and closed‑loop rollouts for several agents.

By Yukuan Lu, Zaishuo Xia, Weyl Lu, Yubei Chen
arXiv Machine Learning
Aug 17

AgilePE: Autonomous UAV Pursuit-Evasion via Self-Play Reinforcement Learning

arXiv:2608. 14135v1 Announce Type: cross Abstract: Autonomous pursuit-evasion is a fundamental challenge for Unmanned Aerial Vehicles (UAVs), requiring rapid decision-making under tightly coupled dynamics and continuously changing opponent behaviors.

By Wenhao Tang, Tianyang Chen, Zhejun Cui, Boyuan An, Jiayu Chen, Ruize Zhang, Huidong Liu, Tianyue Wu, Qingmin Liao, Fei Gao, Yu Wang, Chao Yu
arXiv AI
Jul 17

SportD: Can VLMs Physically Strategize?

arXiv:2607. 14616v1 Announce Type: new Abstract: Vision--language models have become increasingly capable of interpreting visual scenes, but it remains unclear whether they can use information to make strategically effective decisions.

By Jasin Cekinmez, Addison J. Wu, Haotian Xia, Akshaya Bharadhwaj, Anay Putty, Anirudh Ravishankar, Jaewoong Lee, Jinglin Xiao, Kyumin Andrew Shim, Mishika Ahuja, Nisarga Patil, Leo Liu, Zhuohan Liu, Weining Shen
arXiv AI
Jun 10

Monte Carlo Pass Search: Using Trajectory Generation for 3D Counterfactual Pass Evaluation in Football

arXiv:2606. 11120v1 Announce Type: new Abstract: We recast pass evaluation in football (soccer) as a Monte Carlo Tree Search (MCTS)-like evaluation problem whose components mostly exist in the literature under different names: a value model (possession value), a world model (multi-agent trajectories with ball interactions), and a policy over counterfactual actions (sampling pass variants with noise).

By Andrew Kang, Priya Narasimhan