The paper introduces a method for verifying whether two gameplay replays in Counter‑Strike 2 belong to the same player by extracting a behavioral fingerprint that captures crosshair control, movement‑stop‑fire coordination, economy, combat engagement, and temporal rhythm. Using 1,330 demos and 13,300 observations, the authors train a pairwise model that achieves an ROC AUC of 0.931 and 0.722 recall at 95% precision, with low‑level mechanical habits providing the strongest identity signals. Aggregating multiple demos further improves performance, raising AUC to 0.986 when ten historical demos are considered.
By Xuchen Zhang
Faynt is a family of Transformer policies (10M and 75M parameters) that control all 26 characters in Super Smash Bros. Melee from a single checkpoint. After reinforcement learning, the 10M model wins 98.4% of same‑character games against fourteen specialist and multi‑character releases, and defeats a zero‑delay Slippi‑AI model in all 68 evaluated games. The work details architecture, scaling, hyperparameter transfer, supervised pretraining on 840,000 human replays, post‑training curricula, distillation, and efficient inference, and it releases the weights, benchmark suites, and a platform for automated model tournaments.
By Ali Janati, Nikita Kuzmin, Rohit Swamy, Charles Niu
The paper introduces the concept of perfect aliasing, where a truth probe that aligns truthful reporting with a task’s prescribed action cannot differentiate between the two based solely on its labels. In a binary reporting game, probes fitted on compliant contexts yield identical optimizations, while on rival contexts their labels are complementary, causing their AUROCs to sum to one across 751 cell-layer pairs. By employing randomized codebooks and mixed-context fitting, the authors demonstrate that separating prescribed output symbols from semantic action enables perfect recovery of truth, achieving an AUROC of 1.000 on rival trials for a reward-trained Gemma-2-9B policy, whereas conventional probes perform near chance.
By Dylan Jayabahu
arXiv:2608. 16196v1 Announce Type: new Abstract: Personalized game generation requires inferring a player's abilities and behavioral style from how they play.
By Yifan Lu, Xiaopeng Yuan, Haohan Wang
CHAMP is a cross‑domain matchmaking framework for Multiplayer Online Battle Arena games that tackles cold‑start, distribution shift, and data‑scarcity issues by using hybrid player profiles and a Domain‑Aware Win‑rate Network (DAWN). DAWN learns mode‑conditioned representations through a shared network, achieving 67.73% win‑rate prediction accuracy and improving match balance in large‑scale A/B tests. The system reduces imbalanced matches, notably cutting 5‑minute kill crushing rates by up to 20.73% for lower‑tier players.
By Kai Wang, Ge Fan, Chaoyun Zhang, Yuyang Jiang, Yuze Liu
The paper proposes a scalable, automated method for auditing candidate‑job matching systems for demographic bias. It employs large‑language‑model agents to generate neutral resumes, injects controlled demographic variations, ranks candidates with a fine‑tuned embedding model, and evaluates nine fairness metrics across counterfactual, group‑fairness, and merit‑aware families, producing a composite risk report. Experiments on a small corpus show that single‑score audits miss nuanced issues, underscoring the need for multi‑metric evaluation and LLM‑generated audits as a low‑cost complement to human reviews.
By Sai Yashwant, Shruti Bansal, Anurag Dubey, Samaroha Chatterjee, Satyam Kumar, Shreyash Gupta, Gantala Thulsiram
arXiv:2607. 26061v1 Announce Type: new Abstract: Pre-match tactical decision-making in professional football relies heavily on subjective expert analysis and identity-based scouting systems that cannot generalize to unseen teams.
By Mouad Zemzoumi, Amine Abouaomar
arXiv:2608. 01193v1 Announce Type: cross Abstract: An AI development race creates a multi-agent safety dilemma.
By Phu Hoa Pham, Duy Minh Dao Sy, Trung Kiet Huynh, Phu Quy Nguyen Lam, Chi Nguyen Tran, Minh Trung Le, Phong Hao Le, Dinh Nam Nguyen, Thien Ky Nguyen Dong, Elias Fernandez Domingos, Le Hong Trang, The Anh Han
arXiv:2608. 19760v1 Announce Type: cross Abstract: Audited against causal ground truth from executed replay in a single-agent tool environment (ALFWorld), none of the step-level credit signals used to train LLM agents -- LLM-judge scores, outcome-conditioned logprob ratios, or the policy's own confidence -- identifies which steps causally matter better than chance.
By Haiyue Zhang
arXiv:2607. 19369v1 Announce Type: new Abstract: Hidden-state probes often recover latent labels in imperfect-information sequence models, but this alone does not establish that a model maintains a posterior belief distribution over hidden states.
By Quanhao Li, Qianyu Chen
arXiv:2607. 00190v1 Announce Type: cross Abstract: Recent advances in reinforcement learning have produced superhuman agents across a wide range of competitive games.
By Andrzej Bia{\l}ecki, Adam Mastalerz, Han Zhou
arXiv:2609.36587v1 Announce Type: new
Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a prominent approach for improving language-model performance on reasoning tasks using...
By Yupeng Chang, Wenxuan Zhang, Yuan Wu