The paper proposes a method for verifying whether two gameplay replays in Counter‑Strike 2 belong to the same player by extracting a behavioral fingerprint that captures crosshair control, movement, economy, combat, and rhythm. Using a six‑fold evaluation on two datasets (Perfect and Professional), the pairwise model achieves ROC AUCs of 0.926 and 0.956, with aiming and low‑level mechanics providing the strongest identity signals. Aggregating evidence across multiple historical demos further improves account‑history AUC, reaching 0.982 on Perfect and 0.975 on Professional.
By Xuchen Zhang
Faynt is a family of Transformer policies (10M and 75M parameters) that control all 26 characters in Super Smash Bros. Melee from a single checkpoint. After reinforcement learning, the 10M model wins 98.4% of same‑character games against fourteen specialist and multi‑character releases, and defeats a zero‑delay Slippi‑AI model in all 68 evaluated games. The work details architecture, scaling, hyperparameter transfer, supervised pretraining on 840,000 human replays, post‑training curricula, distillation, and efficient inference, and it releases the weights, benchmark suites, and a platform for automated model tournaments.
By Ali Janati, Nikita Kuzmin, Rohit Swamy, Charles Niu
The paper introduces the concept of perfect aliasing, where a truth probe that aligns truthful reporting with a task’s prescribed action cannot differentiate between the two based solely on its labels. In a binary reporting game, probes fitted on compliant contexts yield identical optimizations, while on rival contexts their labels are complementary, causing their AUROCs to sum to one across 751 cell-layer pairs. By employing randomized codebooks and mixed-context fitting, the authors demonstrate that separating prescribed output symbols from semantic action enables perfect recovery of truth, achieving an AUROC of 1.000 on rival trials for a reward-trained Gemma-2-9B policy, whereas conventional probes perform near chance.
By Dylan Jayabahu
CHAMP is a cross‑domain matchmaking framework for Multiplayer Online Battle Arena games that tackles cold‑start, distribution shift, and data‑scarcity issues by using hybrid player profiles and a Domain‑Aware Win‑rate Network (DAWN). DAWN learns mode‑conditioned representations through a shared network, achieving 67.73% win‑rate prediction accuracy and improving match balance in large‑scale A/B tests. The system reduces imbalanced matches, notably cutting 5‑minute kill crushing rates by up to 20.73% for lower‑tier players.
By Kai Wang, Ge Fan, Chaoyun Zhang, Yuyang Jiang, Yuze Liu
arXiv:2608. 19760v1 Announce Type: cross Abstract: Audited against causal ground truth from executed replay in a single-agent tool environment (ALFWorld), none of the step-level credit signals used to train LLM agents -- LLM-judge scores, outcome-conditioned logprob ratios, or the policy's own confidence -- identifies which steps causally matter better than chance.
By Haiyue Zhang
arXiv:2607. 19369v1 Announce Type: new Abstract: Hidden-state probes often recover latent labels in imperfect-information sequence models, but this alone does not establish that a model maintains a posterior belief distribution over hidden states.
By Quanhao Li, Qianyu Chen
arXiv:2608. 01193v1 Announce Type: cross Abstract: An AI development race creates a multi-agent safety dilemma.
By Phu Hoa Pham, Duy Minh Dao Sy, Trung Kiet Huynh, Phu Quy Nguyen Lam, Chi Nguyen Tran, Minh Trung Le, Phong Hao Le, Dinh Nam Nguyen, Thien Ky Nguyen Dong, Elias Fernandez Domingos, Le Hong Trang, The Anh Han
arXiv:2609.36587v1 Announce Type: new
Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a prominent approach for improving language-model performance on reasoning tasks using...
By Yupeng Chang, Wenxuan Zhang, Yuan Wu
arXiv:2610.00385v1 Announce Type: cross
Abstract: Replay selectors often rank cached trajectories by format feedback, confidence, freshness, or response length, although cache-level correctness and d...
By Miaobo Hu, Shuhao Hu, Xiaobo Guo, Xin Wang, Bokun Wang, Tianshu Fu, Daren Zha, Jun Xiao
arXiv:2607. 26061v1 Announce Type: new Abstract: Pre-match tactical decision-making in professional football relies heavily on subjective expert analysis and identity-based scouting systems that cannot generalize to unseen teams.
By Mouad Zemzoumi, Amine Abouaomar
arXiv:2608. 16196v1 Announce Type: new Abstract: Personalized game generation requires inferring a player's abilities and behavioral style from how they play.
By Yifan Lu, Xiaopeng Yuan, Haohan Wang
The paper introduces PAI‑Bench, a benchmark designed to evaluate persistent AI agents on how faithfully they adhere to a versioned identity contract. It separates several dimensions—recall, composition, behavioral enactment, resistance, persistence, lineage, and role‑conditioned updates—while keeping scoring oracles independent of the target process. Experiments on synthetic profiles show that explicit cues can significantly alter the presence of identity identifiers, revealing prompt‑dependent component selection and sensitivity to startup cues.
By Zhenyu Zhao, Roy Zhao