Hugging Face Trending Papers

DynCur-Geo: Dynamic Curiosity Reward Shaping for Multimodal Active Geo-Localization

DynCur-Geo introduces a dynamic curiosity framework for active geo‑localization, adjusting the intrinsic reward based on the remaining distance to a target. A distance‑aware gate promotes early exploration and transitions the policy toward goal‑directed behavior as the UAV approaches the target, while potential‑based reward shaping provides dense progress guidance. Experiments in multimodal, cross‑scene, disaster‑affected, and long‑range scenarios demonstrate consistent performance gains over existing active geo‑localization baselines.

arXiv Computer Vision
Sep 18

Towards Active Cross-View Object Geo-Localization

The paper introduces Active Cross-View Object Geo-Localization (ActiveGeo), enabling mobile agents to actively select new viewpoints and decide when to stop to improve localization with fewer observations. It proposes the ActiveMoPT framework, which uses a three-stage training process: Multi-View Prompt-Preserving Adaptation, Trajectory-Guided Policy Initialization, and Cost-Aware Policy Refinement with GRPO. The authors also create a zero-shot test set, ActiveGeo-858, and demonstrate that ActiveMoPT outperforms prior methods on MoP-UAV and ActiveGeo-858.

By Shunyu Yao, Xiaohan Zhang, Zhuoran Yang, Haoqi Lai, Qi Ming, Xiaoxi Hu, Hui-Liang Shen, Si-Yuan Cao
arXiv Machine Learning
1d ago

Reward as Observation: Learning Reward-Based Policies for Rapid Adaptation

The paper proposes a reward-based policy that relies only on rewards and actions, enabling zero‑shot transfer between source and target environments with entirely different observation spaces. Experiments on Pointmass, Cartpole, 2D Car Racing, and the Stretch robot in Habitat‑Sim show that the policy can adapt to new visual styles or 3D renderings without additional samples. Additionally, the reward policy can guide the training of an observation‑based policy in the target environment.

By Morgan Byrd, Maks Sorokin, Robert Wright, Sehoon Ha
Hugging Face Trending Papers
Aug 6

Uncertainty-Aware World Model for Aerial Image-Goal Navigation

Aerial image-goal navigation requires an unmanned aerial vehicle (UAV) to reach a target location specified by a goal image. Existing world-model-based methods rank candidate trajectories using predicted futures, but typically rely on only one or a few point predictions, which is inadequate for large-scale outdoor environments with substantial future-state uncertainty.

arXiv AI
Jun 3

AirDreamer: Generalist Drone Navigation with World Models

arXiv:2606. 03252v1 Announce Type: cross Abstract: Navigating a drone in unseen and cluttered environments requires reliable generalization to unseen scene layouts and understanding of environmental structure relative to the robot's capabilities.

By Zian Liu, Andong Yang, Chunkai Yang, Ruidong An, Chao Gao, Guyue Zhou
arXiv AI
Aug 19

Integrating Novelty and Surprise for Experience Prioritization and Exploration in Image-Based Reinforcement Learning

The paper proposes Novelty and Surprise Prioritized Experience Replay (NSPER) for image-based reinforcement learning, combining novelty to highlight underrepresented states and surprise to reveal gaps in the agent’s knowledge. An extended version, NSPER+R, also uses these signals as intrinsic rewards to enhance both replay quality and exploration. Experiments on DeepMind Control Suite tasks demonstrate that NSPER and NSPER+R accelerate training and improve convergence compared to existing methods.

By Hoda Yamani, Henry Williams, Bruce A. MacDonald
Hugging Face Trending Papers
Aug 18

Integrating Novelty and Surprise for Experience Prioritization and Exploration in Image-Based Reinforcement Learning

The paper tackles sample efficiency in image-based reinforcement learning by combining novelty and surprise signals to prioritize experiences. It proposes Novelty and Surprise Prioritized Experience Replay (NSPER) and an extended version, NSPER+R, which also uses these signals as intrinsic rewards. Experiments on DeepMind Control Suite tasks demonstrate that both methods accelerate training and improve convergence compared to existing techniques.