LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent
arXiv:2604. 17931v3 Announce Type: replace Abstract: Reinforcement Learning (RL) has emerged as a powerful training paradigm for LLM-based agents.
PhantomEnvironments is a framework that trains large language model agents in synthetic, rule‑generated fictional worlds. By creating multi‑turn reinforcement learning environments where agents search templated articles to answer multi‑hop questions, the approach eliminates the need for costly human data or hallucinated LLM‑generated settings. Agents trained in these zero‑cost, purely rule‑based worlds transfer effectively to real‑world multi‑hop search benchmarks, often surpassing models trained on real data, and demonstrate scalable search behavior that grows linearly with question difficulty.
arXiv:2604. 17931v3 Announce Type: replace Abstract: Reinforcement Learning (RL) has emerged as a powerful training paradigm for LLM-based agents.
arXiv:2601. 21754v3 Announce Type: replace Abstract: While Large Language Models (LLMs) excel in language-based agentic tasks, their applicability to unseen, nonlinguistic environments (e.
arXiv:2605.14211v4 Announce Type: replace Abstract: Long-horizon visuomotor tasks remain a fundamental challenge in AI, as current methods rely on hand-engineered rewards or action-labeled demonstrat...
arXiv:2601.13247v2 Announce Type: replace-cross Abstract: Current Large Language Models (LLMs) exhibit a critical modal disconnect: they possess vast semantic knowledge but lack the procedural ground...
arXiv:2608. 04934v1 Announce Type: cross Abstract: Training LLM agents commonly relies on supervised fine-tuning from expert trajectories or online reinforcement learning over human-specified tasks with handcrafted verifiers.
arXiv:2606. 17680v1 Announce Type: new Abstract: Reinforcement learning (RL) has emerged as a powerful paradigm for training Large Language Models (LLMs) as agents.
arXiv:2609.05576v1 Announce Type: new Abstract: The paradigm of LLMs has rapidly shifted from passive language interfaces to autonomous Claw-like agents that execute long-horizon tasks across statefu...
The paper introduces the Agentic Compositional Generalization hypothesis, suggesting that reinforcement learning (RL) primarily refines high‑level decision‑making behaviors that orchestrate pre‑trained low‑level skills, rather than teaching new domain‑specific skills from scratch. It proposes River, a training recipe that enhances reward quality by filtering low‑quality synthetic environments and adding process‑level behavior regularization. Using River, RL‑trained agents outperform other open‑source 8B models on four terminal‑agent benchmarks, achieving significant gains with fewer than 30% of the training environments.
arXiv:2609.07575v1 Announce Type: cross Abstract: This work introduces an alternative view of efficient exploration and studies its theoretical and empirical implications in the absence of extrinsic...
arXiv:2607. 16204v1 Announce Type: new Abstract: Recent growth in reinforcement learning (RL) has surfaced a need for diverse, specialized training environments.
arXiv:2609.15364v1 Announce Type: new Abstract: Digital agents must often adapt to new environments whose interfaces, tools, and failure modes are not fully captured by pretrained models. We introduc...
MineExplorer is a benchmark designed to assess the open‑world exploration abilities of multimodal large language models (MLLMs) in Minecraft. It filters out tasks that rely heavily on Minecraft‑specific knowledge, organizes tasks into ReAct‑style capabilities, and composes atomic tasks into implicit multi‑hop challenges. A multi‑agent synthesis workflow creates reliable task graphs, sandbox scenes, and rule‑based milestone evaluators, and human evaluation confirms its superiority over a single‑agent baseline. Experiments show that while advanced MLLMs can handle many single‑hop tasks, they struggle with longer trajectories that require coordinating hidden prerequisites, and larger models or different thinking modes do not consistently improve performance.