AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents
arXiv:2607. 02255v1 Announce Type: new Abstract: Memory for a long-horizon LLM agent is a contract about what each future decision is allowed to see.
arXiv:2607. 29577v1 Announce Type: new Abstract: Games and simulators make valuable benchmarks by turning decisions into measurable outcomes, but many current suites under-test rules-rich tactical reasoning: the ability to choose well when geometry, timing, resources, objectives, and rule interactions all matter at once.
arXiv:2607. 02255v1 Announce Type: new Abstract: Memory for a long-horizon LLM agent is a contract about what each future decision is allowed to see.
arXiv:2606. 18950v1 Announce Type: new Abstract: Modern Vision-Language Models (VLMs) often struggle with strategic reasoning, i.
arXiv:2609.21562v1 Announce Type: cross Abstract: Coding agents can modify and test code across large software projects. Game development is a domain where agents must implement gameplay rules. A gam...
arXiv:2608. 00832v1 Announce Type: new Abstract: Structured plan-generation agents are often evaluated as if a plan has quality in isolation, yet many realistic planning tasks require asking how a candidate behaves when another agent can search for responses.
CivBench is an open‑source benchmark that evaluates language‑model agents in the long‑horizon, tool‑mediated game Civilization VI using the Model Context Protocol (MCP). Each episode lasts over 300 turns, generating thousands of tool calls across a 76‑tool action space, and includes a narration layer that translates visual game state into structured text. The study characterises agent behaviour across four model families, introducing Proactive Monitoring Rate (PMR) and RAG@10 as interface‑level metrics, and finds that agents often under‑monitor strategic state and fail to execute near‑term commitments despite tool access and explicit guidance.
arXiv:2608. 09902v1 Announce Type: new Abstract: We introduce the Dark Souls Learning Environment (DSLE), a containerized platform that presents all 22 boss encounters of Dark Souls: Remastered as game-playing agent benchmarks through a Gymnasium-style interface.
Faynt is a family of Transformer policies (10M and 75M parameters) that control all 26 characters in Super Smash Bros. Melee from a single checkpoint. After reinforcement learning, the 10M model wins 98.4% of same‑character games against fourteen specialist and multi‑character releases, and defeats a zero‑delay Slippi‑AI model in all 68 evaluated games. The work details architecture, scaling, hyperparameter transfer, supervised pretraining on 840,000 human replays, post‑training curricula, distillation, and efficient inference, and it releases the weights, benchmark suites, and a platform for automated model tournaments.
arXiv:2609.38881v1 Announce Type: new Abstract: Real-time strategy (RTS) games require agents to coordinate economic development, production and construction, base defense, unit organization, and att...
The paper introduces WebMRE, an offline benchmark comprising 541 tasks and 5,293 steps extracted from WebArena trajectories, designed to provide deterministic scoring for web agents without live environments. It enables the first systematic study of how guide sentences and grounded actions reinforce each other, showing that jointly decoding a guide improves element selection accuracy and that the guide acts as a causal instruction channel. The authors fine‑tune models that outperform leading zero‑shot baselines on all offline metrics.
arXiv:2605.11467v2 Announce Type: replace-cross Abstract: Reasoning models post-hoc rationalize answers they have already committed to internally, producing chains of *reasoning theater*: deliberativ...
arXiv:2605. 11030v2 Announce Type: replace-cross Abstract: Closed-loop tool-using agents are increasingly evaluated in executable web, code, and micro-task environments, but benchmark reports often conflate workloads, action-generating drivers, and the evidence admitted for systems-facing claims.
arXiv:2606. 01667v1 Announce Type: new Abstract: Test-time scaling has become a major way to improve large language model reasoning, but its orchestration has remained designer-engineered: a fixed sample budget, a fixed refinement loop, a fixed scoring rule, or a fixed search policy decides how compute is spent, leaving the model in charge of solving but not of orchestration.