The Rise of Verbal Reinforcement Learning
arXiv:2609.01597v1 Announce Type: cross Abstract: Natural language is emerging as a primary feedback channel for improving language agents, capable of conveying intent, preferences, and causal struct...
The paper investigates how human observers interpret the learning processes of reinforcement learning (RL) agents. Using a novel observation-based paradigm, the authors conducted two experiments: an exploratory interview study with nine participants that identified four core themes—Agent Goals, Knowledge, Decision Making, and Learning Mechanisms—and a confirmatory study with 34 participants that applied the paradigm across navigation and manipulation tasks and two RL algorithms. Analyses of 816 responses validated the paradigm’s reliability and refined the thematic framework, showing how these themes evolve over time and interrelate.
arXiv:2609.01597v1 Announce Type: cross Abstract: Natural language is emerging as a primary feedback channel for improving language agents, capable of conveying intent, preferences, and causal struct...
CoLearn is an interactive, agentic tutoring system that learns about each learner through a persistent memory of mastery and misconceptions, updated with a Bayesian Knowledge Tracing model that uses a large language model as an observation function. It generates personalized questions targeting the learner’s weakest topics and recurring misconceptions, and provides a live evidence view for progress visualization and blind A/B comparison. In blind A/B tests, learners preferred questions conditioned on this memory 68‑69% of the time, and simulations show the agent’s belief converges toward the learner’s true mastery.
The paper explores Retrospection-Only Fine-Tuning (ROFT), a method where a language-model agent improves its behavior by generating and training on explanations of its own experiences, without external teachers or reward signals. In software‑engineering tasks with Qwen3.5‑4B, ROFT achieves comparable or better solve rates than GRPO while requiring fewer updates and training time, and can learn from failures alone. Behavioral analysis shows ROFT indirectly assigns credit to actions and can produce shorter, more direct solutions when prompted to focus on direct solutions.
arXiv:2606. 15306v1 Announce Type: cross Abstract: We envision continually learning agentic systems that become more useful over time: as they encounter sequences of related tasks, they should infer the hidden structure shared across those tasks and use it to improve future decisions.
Teach-and-Grow Learning (TGL) is an agent-centered architecture that transforms a few successful demonstrations into reusable Skill Blocks, enabling a robot to compose, execute, and revise behaviors in new scenes without task-specific policy retraining. The system maintains a Skill Library and structured Experience Memory to capture successes, failures, and repairs, allowing persistent reuse and agent-directed adaptation. Evaluation on the LIBERO benchmark shows state-of-the-art performance, and the authors propose a scaling-law hypothesis suggesting that accumulated reusable experience reduces future-task error and teaching demand following a power-law trend.
arXiv:2607. 28638v1 Announce Type: cross Abstract: As large language model (LLM) agents increasingly learn from experience, they primarily rely on trajectory-level reflection to extract insights.
arXiv:2606. 17680v1 Announce Type: new Abstract: Reinforcement learning (RL) has emerged as a powerful paradigm for training Large Language Models (LLMs) as agents.
arXiv:2609.19138v1 Announce Type: new Abstract: Enabling robots to adapt to unfamiliar environments as readily as humans remains a moonshot goal of embodied AI. No finite collection of demonstrations...
arXiv:2410.23912v3 Announce Type: replace-cross Abstract: The reasoning abilities of large language models (LLMs) have improved with chain-of-thought (CoT) prompting, allowing models to solve complex...
arXiv:2608. 04934v1 Announce Type: cross Abstract: Training LLM agents commonly relies on supervised fine-tuning from expert trajectories or online reinforcement learning over human-specified tasks with handcrafted verifiers.
arXiv:2606. 01672v1 Announce Type: new Abstract: Reinforcement learning has enabled the acquisition of impressive robotic skills, but typically requires hand-crafted reward functions that are slow to design and difficult to align with human intentions.
arXiv:2606. 27136v1 Announce Type: new Abstract: For LLM agents in multi-step interactive environments, a key challenge is to make effective use of accumulated interaction experience.