arXiv:2608. 05206v1 Announce Type: new Abstract: Otter is a 15.
By Tarun Kumar S
arXiv:2606. 25176v2 Announce Type: replace Abstract: Chess engines have evolved from search-based systems optimized solely for strength to neural policies capable of modeling human decisions across much of the rating spectrum.
By Jason Carlson
arXiv:2606. 26267v1 Announce Type: new Abstract: Rating systems such as Elo serve as the gold standard for matchmaking in competitive chess.
By Tianyuan Zhou, Zhizheng Fu, Tianming Yang
arXiv:2608. 03416v1 Announce Type: new Abstract: Large language models (LLMs) are now regularly asked to forecast real-world events, but comparisons are often difficult because models receive different information, use different tools, and are evaluated under different rules.
By Jonaid Shianifar, Iias Faiud
arXiv:2607. 17765v1 Announce Type: cross Abstract: We introduce WC2026-Agents, a benchmark and dataset for evaluating large language models (LLMs) as autonomous forecasting agents on real, future events.
By Jiacheng Ding, Cong Guo, Jason Xu
Faynt is a family of Transformer policies (10M and 75M parameters) that control all 26 characters in Super Smash Bros. Melee from a single checkpoint. After reinforcement learning, the 10M model wins 98.4% of same‑character games against fourteen specialist and multi‑character releases, and defeats a zero‑delay Slippi‑AI model in all 68 evaluated games. The work details architecture, scaling, hyperparameter transfer, supervised pretraining on 840,000 human replays, post‑training curricula, distillation, and efficient inference, and it releases the weights, benchmark suites, and a platform for automated model tournaments.
By Ali Janati, Nikita Kuzmin, Rohit Swamy, Charles Niu