Can In-Context Learning Support Intrinsic Curiosity?
arXiv:2606. 19476v1 Announce Type: cross Abstract: Effective machine learning depends not only on how we model data, but also on what data we choose to collect.
arXiv:2606. 19476v1 Announce Type: cross Abstract: Effective machine learning depends not only on how we model data, but also on what data we choose to collect.
Effective machine learning depends not only on how we model data, but also on what data we choose to collect. While large sequence models have revolutionized data modeling, the problem of automated data selection, or "intrinsic curiosity", remains a significant challenge.
We’ve developed Random Network Distillation (RND), a prediction-based method for encouraging reinforcement learning agents to explore their environments through curiosity, which for the first time exceeds average human performance on Montezuma’s Revenge.
The paper introduces Gradient‑Momentum Coupling (GMC), a method that quantifies learning progress by measuring how strongly a sample influences changes in the parameter space, using the normalized absolute product of its gradient and the momentum of previous gradients. GMC filters out noise by accumulating consistent directions of change while canceling random fluctuations, leading to a more uniform prioritization across tasks with varying noise levels and better ranking of learnable tasks by improvement speed. Experiments on MiniGrid MultiRoom tasks show that replacing prediction error with GMC in the Intrinsic Curiosity Module restores exploration capabilities that were lost to unpredictable observations.
arXiv:2604. 18701v3 Announce Type: replace-cross Abstract: Local prediction-error-based curiosity rewards focus on the current transition without considering the world model's cumulative prediction error across all visited transitions.
arXiv:2503. 14833v2 Announce Type: replace-cross Abstract: One of the bottlenecks in robotic intelligence is the instability of neural network models.
arXiv:2609.07575v1 Announce Type: cross Abstract: This work introduces an alternative view of efficient exploration and studies its theoretical and empirical implications in the absence of extrinsic...
arXiv:2606. 12386v1 Announce Type: cross Abstract: Advancing scientific understanding through mechanistic modeling requires posing the right experimental questions to yield maximally informative data.
arXiv:2607. 17981v1 Announce Type: new Abstract: Representation learning has enabled classical exploration strategies to be extended to deep Reinforcement Learning (RL), but often makes algorithms more complex and theoretical guarantees harder to establish.
DynCur-Geo introduces a dynamic curiosity framework for active geo‑localization, adjusting the intrinsic reward based on the remaining distance to a target. A distance‑aware gate promotes early exploration and transitions the policy toward goal‑directed behavior as the UAV approaches the target, while potential‑based reward shaping provides dense progress guidance. Experiments in multimodal, cross‑scene, disaster‑affected, and long‑range scenarios demonstrate consistent performance gains over existing active geo‑localization baselines.
arXiv:2602. 02244v3 Announce Type: replace Abstract: The standard post-training recipe for large reasoning models, supervised fine-tuning followed by reinforcement learning (SFT-then-RL), may limit the benefits of the RL stage: while SFT imitates expert demonstrations, it often causes overconfidence and reduces generation diversity, leaving RL with a narrowed solution space to explore.