Hugging Face Trending Papers

DynCur-Geo: Dynamic Curiosity Reward Shaping for Multimodal Active Geo-Localization

Read the original on Hugging Face Trending Papers →

DynCur-Geo introduces a dynamic curiosity framework for active geo‑localization, adjusting the intrinsic reward based on the remaining distance to a target. A distance‑aware gate promotes early exploration and transitions the policy toward goal‑directed behavior as the UAV approaches the target, while potential‑based reward shaping provides dense progress guidance. Experiments in multimodal, cross‑scene, disaster‑affected, and long‑range scenarios demonstrate consistent performance gains over existing active geo‑localization baselines.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv Computer Vision
Sep 18

Towards Active Cross-View Object Geo-Localization

The paper introduces Active Cross-View Object Geo-Localization (ActiveGeo), enabling mobile agents to actively select new viewpoints and decide when to stop to improve localization with fewer observations. It proposes the ActiveMoPT framework, which uses a three-stage training process: Multi-View Prompt-Preserving Adaptation, Trajectory-Guided Policy Initialization, and Cost-Aware Policy Refinement with GRPO. The authors also create a zero-shot test set, ActiveGeo-858, and demonstrate that ActiveMoPT outperforms prior methods on MoP-UAV and ActiveGeo-858.

By Shunyu Yao, Xiaohan Zhang, Zhuoran Yang, Haoqi Lai, Qi Ming, Xiaoxi Hu, Hui-Liang Shen, Si-Yuan Cao
arXiv Machine Learning
1d ago

Reward as Observation: Learning Reward-Based Policies for Rapid Adaptation

The paper proposes a reward-based policy that relies only on rewards and actions, enabling zero‑shot transfer between source and target environments with entirely different observation spaces. Experiments on Pointmass, Cartpole, 2D Car Racing, and the Stretch robot in Habitat‑Sim show that the policy can adapt to new visual styles or 3D renderings without additional samples. Additionally, the reward policy can guide the training of an observation‑based policy in the target environment.

By Morgan Byrd, Maks Sorokin, Robert Wright, Sehoon Ha