arXiv AI

Engineering Efficient Self-Play Chess: Search, Replay, and Throughput Under Limited Compute

arXiv AI
1d ago

KV-streams for Efficient Compaction in Agentic Reinforcement Learning

arXiv:2609.35750v2 Announce Type: replace-cross Abstract: Scaling the horizon of agentic LLMs is bottlenecked by the need to fit ever longer context traces in GPU memory. Context compaction has been...

By Emiliano Penaloza, Dane Malenfant, Dheeraj Vattikonda, Roger Creus Castanyer, Siddarth Venkatraman, Abhay Puri, Jonathan Light, Matthew James Sargent, Augustine N. Mavor-Parker, Massimo Caccia, Lucas Caccia, Glen Berseth, Esmeralda S. Whitammer, Alessandro Sordoni, Minseon Kim, Marc-Alexandre C\^ot\'e, Laurent Charlin, Guillaume Lajoie
arXiv Computation and Language
Sep 15

Dream-RSI: Recursive Self-Improvement through Evolving Worlds

arXiv:2609.14858v1 Announce Type: new Abstract: Recursive self-improvement is becoming increasingly vital for autonomous AI agents, where progress hinges on discovering high-value solutions across co...

By Tong Zheng, Xidong Wu, Zheng Zhang, Zhankui He, Chaoyi Zhang, Benjamin Coleman, Ruoqiao Wei, Di Bai, Haolin Liu, Rui Liu, Xue Wang, Yue Zhuan, Wang-Cheng Kang, Renkai Xiang, Heng Huang, Xinwu Cheng, Yunsong Guo
arXiv Machine Learning
Sep 14

Code-to-Harness: Distilling Black-Box Optimizers from Self-Play

The paper investigates whether an agent can learn a numerical search strategy through executable practice and then encode that strategy as text. By repeatedly writing and evaluating optimizer programs, the agent distills a 197‑word text called Harness A, which significantly reduces regret for Gemini Flash and other language‑model executors, matching the performance of classical Gaussian‑process Bayesian optimization. An independent replication produced a different but equally effective text, Harness B, and the framework also achieved the lowest regret on a sealed YouTube reward‑tuning benchmark.

By Yi Wu, Zheng Ren, Zhiyu Hu, Haochen Wang, Daryl Chang, Li Wei, Ting Wang, Zhen Li, Pooja Gupta, Nitin Jindal, Lukasz Heldt
arXiv AI
Jun 2

MindGames Arena Generalization Track: In2AI Solution with Delayed Per-Step Reward Attribution

arXiv:2606. 00017v1 Announce Type: new Abstract: Training language model agents for multi-agent strategic interaction presents a core difficulty: the quality of any action may depend on future events that never materialize, on moves that violate game rules, or on decisions made by other players.

By Aliaksei Korshuk, Alexander Buyantuev, Ilya Makarov
arXiv Machine Learning
Aug 19

Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements

Agentic ESOpt proposes using evolution strategies (ES) instead of reinforcement learning to fine‑tune large language‑model agents for long‑horizon tasks. ES offers model scalability, flexibility, and better long‑horizon credit assignment, enabling full‑parameter optimization with minimal GPU memory. The framework samples parameter perturbations, evaluates agents with rewards, and updates online, achieving notable performance gains on WebArena‑Lite and in test‑time prompt‑parameter co‑evolution.

By Zhi Zheng, Rongsheng Chen, Yunpeng Ba, Zhenkun Wang, Yee Whye Teh, Wee Sun Lee