Can In-Context Learning Support Intrinsic Curiosity?
arXiv:2606. 19476v1 Announce Type: cross Abstract: Effective machine learning depends not only on how we model data, but also on what data we choose to collect.
arXiv:2604. 18701v3 Announce Type: replace-cross Abstract: Local prediction-error-based curiosity rewards focus on the current transition without considering the world model's cumulative prediction error across all visited transitions.
arXiv:2606. 19476v1 Announce Type: cross Abstract: Effective machine learning depends not only on how we model data, but also on what data we choose to collect.
Effective machine learning depends not only on how we model data, but also on what data we choose to collect. While large sequence models have revolutionized data modeling, the problem of automated data selection, or "intrinsic curiosity", remains a significant challenge.
arXiv:2510.15047v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) as agents often fail to improve in new environments. We identify and characterize a failure mode we call explora...
The paper introduces Feedback‑Enriched Environments (FEEs) as a new approach to training large language models as autonomous agents for long‑horizon tasks. By shifting from action guidance to observation enrichment during later stages of exploration, FEEs improve performance across SciWorld and BFCL benchmarks with various Qwen3 model scales and RL algorithms. The study shows that FEEs stabilize training, promote proactive exploration, embed environmental guidance into policy weights, and highlight intra‑group feedback consistency as key for stable optimization.
arXiv:2503. 14833v2 Announce Type: replace-cross Abstract: One of the bottlenecks in robotic intelligence is the instability of neural network models.
arXiv:2607. 16858v1 Announce Type: cross Abstract: Across environments with mixed sources of uncertainty, unsupervised reinforcement learning requires intrinsic motivation that does not precommit to a particular direction of surprise.
arXiv:2609.05650v1 Announce Type: new Abstract: We propose a reinforcement learning framework in which exploration is driven by intrinsic curiosity, designed for scenarios where environments are non-...
arXiv:2609.38329v1 Announce Type: new Abstract: Group-relative RL methods such as Flow-GRPO post-train image generators by exploring with isotropic Gaussian noise added at every denoising step. This...
arXiv:2606. 30789v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) has become a standard tool for improving the reasoning ability of large language models, yet its training dynamics are still described empirically: reward trajectories are fit with low-parameter functional forms whose constants carry no mechanistic meaning, and hyperparameter choices remain a matter of trial and error.
arXiv:2602. 02244v3 Announce Type: replace Abstract: The standard post-training recipe for large reasoning models, supervised fine-tuning followed by reinforcement learning (SFT-then-RL), may limit the benefits of the RL stage: while SFT imitates expert demonstrations, it often causes overconfidence and reduces generation diversity, leaving RL with a narrowed solution space to explore.
arXiv:2606. 08815v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has emerged as a powerful paradigm for eliciting long-chain reasoning in large language models.
The paper introduces ELEMENT, a framework that combines episodic and lifelong entropy maximization to drive reward-free exploration in reinforcement learning. It addresses two key limitations of existing entropy-based methods: the vanishing intrinsic reward after a state is visited and the computational cost of estimating entropy over large datasets. ELEMENT achieves this by deriving an average episodic state entropy reward and employing a k‑NN graph‑based estimator for lifelong entropy, leading to superior state coverage and unsupervised pre‑training performance compared to current baselines.