arXiv AI By Xuchen Zhang

Account Consistency from Gameplay Traces: Same-Player Verification in Counter-Strike 2

Read the original on arXiv AI →

The paper proposes a method for verifying whether two gameplay replays in Counter‑Strike 2 belong to the same player by extracting a behavioral fingerprint that captures crosshair control, movement, economy, combat, and rhythm. Using a six‑fold evaluation on two datasets (Perfect and Professional), the pairwise model achieves ROC AUCs of 0.926 and 0.956, with aiming and low‑level mechanics providing the strongest identity signals. Aggregating evidence across multiple historical demos further improves account‑history AUC, reaching 0.982 on Perfect and 0.975 on Professional.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Aug 27

Same-Player Verification for Account Consistency in Counter-Strike 2

The paper introduces a method for verifying whether two gameplay replays in Counter‑Strike 2 belong to the same player by extracting a behavioral fingerprint that captures crosshair control, movement‑stop‑fire coordination, economy, combat engagement, and temporal rhythm. Using 1,330 demos and 13,300 observations, the authors train a pairwise model that achieves an ROC AUC of 0.931 and 0.722 recall at 95% precision, with low‑level mechanical habits providing the strongest identity signals. Aggregating multiple demos further improves performance, raising AUC to 0.986 when ten historical demos are considered.

By Xuchen Zhang
arXiv Machine Learning
1d ago

Faynt: Scaling and Optimizing Policies for Competitive Melee

Faynt is a family of Transformer policies (10M and 75M parameters) that control all 26 characters in Super Smash Bros. Melee from a single checkpoint. After reinforcement learning, the 10M model wins 98.4% of same‑character games against fourteen specialist and multi‑character releases, and defeats a zero‑delay Slippi‑AI model in all 68 evaluated games. The work details architecture, scaling, hyperparameter transfer, supervised pretraining on 840,000 human replays, post‑training curricula, distillation, and efficient inference, and it releases the weights, benchmark suites, and a platform for automated model tournaments.

By Ali Janati, Nikita Kuzmin, Rohit Swamy, Charles Niu
arXiv Machine Learning
Sep 11

The Truth Was Never Gone: Perfect Aliasing in Compliant-Context Truth Probes

The paper introduces the concept of perfect aliasing, where a truth probe that aligns truthful reporting with a task’s prescribed action cannot differentiate between the two based solely on its labels. In a binary reporting game, probes fitted on compliant contexts yield identical optimizations, while on rival contexts their labels are complementary, causing their AUROCs to sum to one across 751 cell-layer pairs. By employing randomized codebooks and mixed-context fitting, the authors demonstrate that separating prescribed output symbols from semantic action enables perfect recovery of truth, achieving an AUROC of 1.000 on rival trials for a reward-trained Gemma-2-9B policy, whereas conventional probes perform near chance.

By Dylan Jayabahu
arXiv AI
Sep 7

CHAMP: Cross-domain Hybrid Architecture for Matchmaking and Prediction in Online Multi-Player Games

CHAMP is a cross‑domain matchmaking framework for Multiplayer Online Battle Arena games that tackles cold‑start, distribution shift, and data‑scarcity issues by using hybrid player profiles and a Domain‑Aware Win‑rate Network (DAWN). DAWN learns mode‑conditioned representations through a shared network, achieving 67.73% win‑rate prediction accuracy and improving match balance in large‑scale A/B tests. The system reduces imbalanced matches, notably cutting 5‑minute kill crushing rates by up to 20.73% for lower‑tier players.

By Kai Wang, Ge Fan, Chaoyun Zhang, Yuyang Jiang, Yuze Liu
arXiv AI
Aug 28

Counterfactual Bias Testing for Application Tracking System

The paper proposes a scalable, automated method for auditing candidate‑job matching systems for demographic bias. It employs large‑language‑model agents to generate neutral resumes, injects controlled demographic variations, ranks candidates with a fine‑tuned embedding model, and evaluates nine fairness metrics across counterfactual, group‑fairness, and merit‑aware families, producing a composite risk report. Experiments on a small corpus show that single‑score audits miss nuanced issues, underscoring the need for multi‑metric evaluation and LLM‑generated audits as a low‑cost complement to human reviews.

By Sai Yashwant, Shruti Bansal, Anurag Dubey, Samaroha Chatterjee, Satyam Kumar, Shreyash Gupta, Gantala Thulsiram