The paper introduces Regret-Weighted Payoff Sampling (RWPS), a budgeted estimator that selectively simulates only payoff-matrix cells relevant to a Nash equilibrium and uses a surrogate model for the remaining entries. RWPS provides an instance-dependent error bound weighted by the opponent’s equilibrium mixture and a coverage result guaranteeing that, once the deviation-relevant set is simulated, surrogate error does not affect either player’s regret. Experiments on three 21×21 general-sum games, including an asymmetric Colonel Blotto, show that RWPS achieves four to six times tighter bounds than previous methods and outperforms other sampling strategies on the CyGym and ANSG cyber simulators at low budgets.
By Michael Lanier, David Farmer, Yevgeniy Vorobeychik
FM‑Bench is a new benchmark that tests large language model agents on long‑horizon decision making by having them run a football club for 20 in‑game years. The agent must manage a squad, trade players, negotiate contracts, invest in facilities and youth, set lineups, and respond to a board that can fire it, all using 26 tools and roughly 340–400 decision stops, with a deterministic engine producing a final score without human or LLM judges. The benchmark includes a solo track where each of 15 frontier models competes against a frozen scripted world, and an Arena track where the same models plus a scripted anchor share one 20‑year world, allowing the first head‑to‑head evaluation at this scale.
whyItMatters":"FM‑Bench provides a rigorous, large‑scale test of sustained, cumulative decision‑making in language‑model agents, revealing that managerial strategy—not computational scale or vendor—drives performance over long horizons."
By Tianyou Wang, Chongyang Gao, Kezhen Chen, Chen Dong, Yinghao He, Donghan Li, Wangcheng Xu, Hongjiu Zhang, Chi Li
FM‑Bench is a new benchmark that tests large language model agents on long‑horizon decision‑making by having them run a football club for 20 in‑game years. The agent must manage a squad, trade players, negotiate contracts, invest in facilities, set lineups, and respond to a board that can fire it, all while a deterministic engine aggregates the outcomes into a final score without human judgment. The benchmark evaluates six behavioral capabilities and compares 15 frontier models in solo and arena tracks, revealing that managerial behavior—not computational scale—drives performance.
arXiv:2607. 02255v1 Announce Type: new Abstract: Memory for a long-horizon LLM agent is a contract about what each future decision is allowed to see.
By Xiangchen Cheng, Yunwei Jiang, Jianwen Sun, Zizhen Li, Chuanhao Li, Xiangcheng Cao, Yihao Liu, Fanrui Zhang, Li Jin, Kaipeng Zhang
arXiv:2606. 29169v1 Announce Type: cross Abstract: Many important games have more than two players and imperfect information.
By Sam Ganzfried
arXiv:2607. 10960v1 Announce Type: new Abstract: Trader-facing dynamic fees are increasingly proposed for automated market makers (AMMs), but historical data do not identify how order flow would respond: trader-facing fees do not vary, trader types are latent, and a replayed tape is not a sequential decision environment.
By Wen-Ting Wang