arXiv AI

Beyond Tracking or Shortcut: Composition-Bounded Predictive States in Poker Autoregressive Models

arXiv:2607. 19369v1 Announce Type: new Abstract: Hidden-state probes often recover latent labels in imperfect-information sequence models, but this alone does not establish that a model maintains a posterior belief distribution over hidden states.

arXiv AI
Aug 26

Confident at the moment of action: belief miscalibration in LLM play under hidden information

The paper investigates whether large language models (LLMs) correctly gauge their confidence when acting in a hidden‑information chess variant. In experiments where the location of a hidden royal piece is repeatedly relocated, the models’ stated probabilities about the piece’s position were almost never accurate at high confidence levels, with a calibration deficit concentrated in those high‑confidence events. Across multiple model configurations and providers, the same pattern emerged, and conventional evaluation metrics such as legality, cost, latency, and completion rate were found to be uncorrelated with belief quality, yet a model could still win the game despite poor confidence estimates.

By Bhushan Kashinath Joshi
arXiv AI
Sep 25

AD-WM: Action-Discriminative World Models for Counterfactual Model Predictive Control

AD-WM is a new action‑discriminative joint‑embedding world model designed for counterfactual model predictive control. It augments residual latent dynamics with action‑recovery regularization based on inverse dynamics and conditional mutual information, while discarding auxiliary heads at test time so that MPC remains unchanged. Experiments on OGBench‑Cube and other simulation environments show substantial gains in hard‑start success and mean success, and zero‑shot transfer to a Franka robot improves pick‑and‑place success from 42.2% to 71.1%.

By Jiabin Qiu, Zixuan Chen, Hongye Cao, Jieqi Shi, Jing Huo, Yang Gao
arXiv AI
Sep 23

TimeLitmus: A Diagnostic Benchmark for Cross-Modal Understanding and Explanation Faithfulness in Event-Conditioned Time-Series Prediction

arXiv:2609.24677v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to make predictions from numerical time-series histories and textual events. Yet accuracy alone cann...

By Jie Gong, Maowei Jiang, Zhiwei Liu, Yankai Chen, Guojun Xiong, Xue Liu, Min Peng, Qianqian Xie, Sophia Ananiadou
arXiv Machine Learning
Aug 27

Same-Player Verification for Account Consistency in Counter-Strike 2

The paper introduces a method for verifying whether two gameplay replays in Counter‑Strike 2 belong to the same player by extracting a behavioral fingerprint that captures crosshair control, movement‑stop‑fire coordination, economy, combat engagement, and temporal rhythm. Using 1,330 demos and 13,300 observations, the authors train a pairwise model that achieves an ROC AUC of 0.931 and 0.722 recall at 95% precision, with low‑level mechanical habits providing the strongest identity signals. Aggregating multiple demos further improves performance, raising AUC to 0.986 when ten historical demos are considered.

By Xuchen Zhang
arXiv AI
Aug 28

Account Consistency from Gameplay Traces: Same-Player Verification in Counter-Strike 2

The paper proposes a method for verifying whether two gameplay replays in Counter‑Strike 2 belong to the same player by extracting a behavioral fingerprint that captures crosshair control, movement, economy, combat, and rhythm. Using a six‑fold evaluation on two datasets (Perfect and Professional), the pairwise model achieves ROC AUCs of 0.926 and 0.956, with aiming and low‑level mechanics providing the strongest identity signals. Aggregating evidence across multiple historical demos further improves account‑history AUC, reaching 0.982 on Perfect and 0.975 on Professional.

By Xuchen Zhang