arXiv Machine Learning

Endogenous Exploration in Reinforcement Learning with Intrinsic Curiosity

arXiv AI
3d ago

Tactile Curiosity Drives Robot Interaction

The paper introduces TacEx, a tactile‑curiosity framework that guides reinforcement learning agents to explore contact dynamics by focusing epistemic uncertainty on the tactile channel. By anchoring curiosity to touch, robots learn to manipulate and grasp objects without task rewards or demonstrations, generating an interaction‑dense dataset that supports offline pick‑and‑place policy learning. TacEx also enhances vision‑language‑action models through post‑training, improving downstream performance while remaining sample‑efficient.

By Klemens Iten, Alexander Proshkin, Bhavya Sukhija, Stelian Coros, Andreas Krause, Pieter Abbeel, Carmelo Sferrazza
arXiv AI
Sep 10

Environments as Scaffold: Enriching Feedback to Bootstrap Self-Evolving Agents in Long-Horizon Tasks

The paper introduces Feedback‑Enriched Environments (FEEs) as a new approach to training large language models as autonomous agents for long‑horizon tasks. By shifting from action guidance to observation enrichment during later stages of exploration, FEEs improve performance across SciWorld and BFCL benchmarks with various Qwen3 model scales and RL algorithms. The study shows that FEEs stabilize training, promote proactive exploration, embed environmental guidance into policy weights, and highlight intra‑group feedback consistency as key for stable optimization.

By Hongbang Yuan, Zhuoran Jin, Yixin Cao
arXiv AI
Jun 2

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying

arXiv:2606. 00151v1 Announce Type: cross Abstract: In reinforcement learning (RL), agents benefit from exploration only because they repeatedly encounter similar states: trying different actions can improve performance or reduce uncertainty; without such retries, a greedy policy is optimal.

By Soichiro Nishimori, Paavo Parmas, Sotetsu Koyamada, Tadashi Kozuno, Toshinori Kitamura, Shin Ishii, Yutaka Matsuo
arXiv Computation and Language
Sep 10

Why Do LLM Agents Fail in Exploring New Environments? A World-Modeling Perspective

arXiv:2510.15047v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) as agents often fail to improve in new environments. We identify and characterize a failure mode we call explora...

By Shiqi Chen, Tongyao Zhu, Zian Wang, Jinghan Zhang, Kangrui Wang, Ruochen Zhou, Siyang Gao, Teng Xiao, Yee Whye Teh, Junxian He, Manling Li
arXiv AI
Aug 12

Exploration-Driven Personalized Federated Reinforcement Learning via Intrinsic Motivation

arXiv:2608. 10499v1 Announce Type: cross Abstract: Personalized Federated Reinforcement Learning (PFRL) takes a decentralized approach to storing and accessing information based on past experiences while keeping each client's data private during the learning of each client's policy.

By Md Rafid Islam, Rafsan Jany, Zahid Hasan, Ratun Rahman
arXiv Machine Learning
1d ago

When Do Intrinsic Rewards Lead to Exploration?

The paper investigates when intrinsic rewards effectively drive exploration in reinforcement learning. It introduces a formal criterion that evaluates policies based on the counterfactual information they acquire, comparing how well their histories can replace experience from alternative policies. Using a simple environment, the authors show that common intrinsic reward objectives—count-based, prediction-error, empowerment, and information-gain—can lead to Pareto-suboptimal exploration under this criterion, and they propose conditions and a new objective that better align with optimal exploration.

By Scott W. Viteri (Stanford University), Laura Gomezjurado Gonzalez (Stanford University), Clark Barrett (Stanford University)
arXiv Machine Learning
Sep 11

From Connectivity to Rewards: Dense Reward Learning with Directed State Graphs

The paper introduces Graph-Guided Quasimetric Dense Reward (G2QDR), a framework that learns a state connectivity model to predict pairwise connectivity strengths in asymmetric environments. These strengths are converted into scalar auxiliary dense rewards, offering continuous guidance across hierarchical levels. G2QDR can be integrated into any existing Goal-Conditioned Hierarchical Reinforcement Learning architecture and shows empirical performance improvements in sparse reward settings with modest computational cost.

By Shuyuan Zhang, Zihan Wang, Xiao-Wen Chang, Doina Precup