arXiv Machine Learning
Aug 19

Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements

Agentic ESOpt proposes using evolution strategies (ES) instead of reinforcement learning to fine‑tune large language‑model agents for long‑horizon tasks. ES offers model scalability, flexibility, and better long‑horizon credit assignment, enabling full‑parameter optimization with minimal GPU memory. The framework samples parameter perturbations, evaluates agents with rewards, and updates online, achieving notable performance gains on WebArena‑Lite and in test‑time prompt‑parameter co‑evolution.

By Zhi Zheng, Rongsheng Chen, Yunpeng Ba, Zhenkun Wang, Yee Whye Teh, Wee Sun Lee
arXiv AI
Aug 11

Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models

arXiv:2608. 09696v1 Announce Type: new Abstract: Predicting the answer to interventional ``what if'' questions --- the outcome of an action never taken --- requires a \emph{mechanistic}, causal model, not a curve fit; and learning such a model requires \emph{experiments}, because passive data leaves its mechanisms unidentified.

By Kevin Murphy
Hugging Face Trending Papers
Aug 10

Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models

Predicting the answer to interventional ``what if'' questions --- the outcome of an action never taken --- requires a \emph{mechanistic}, causal model, not a curve fit; and learning such a model requires \emph{experiments}, because passive data leaves its mechanisms unidentified. Experiments are expensive, so the central problem is \emph{data efficiency}.

arXiv AI
3d ago

PhantomEnvironments: Training LLM Agents in Fictional Worlds

PhantomEnvironments is a framework that trains large language model agents in synthetic, rule‑generated fictional worlds. By creating multi‑turn reinforcement learning environments where agents search templated articles to answer multi‑hop questions, the approach eliminates the need for costly human data or hallucinated LLM‑generated settings. Agents trained in these zero‑cost, purely rule‑based worlds transfer effectively to real‑world multi‑hop search benchmarks, often surpassing models trained on real data, and demonstrate scalable search behavior that grows linearly with question difficulty.

By Anmol Kabra, Swathi Saravana Selvam, Albert Gong, Chao Wan, Christian Belardi, Dongyoung Go, Katie Z. Luo, Kilian Q. Weinberger