arXiv AI

Arm2Air: Cross-Embodiment Skeleton Transfer for 3D Relay Formation

arXiv AI
Sep 7

Continual Field-Adaptive Models (CFAMs) for Post-Deployment Physical AI

arXiv:2609. 04552v1 Announce Type: cross Abstract: Unattended interactive autonomy - machines that step into danger in place of humans and complete tasks with human tools - remains a missing capability in mission-critical operations.

By Amarjot Singh, Tanmay R. Pancholi, Jainam Kothari, Shrirang Mahajan, Ketan Bansal, Zackory Erickson, Giuseppe Loianno, Alexandre M. Bayen, Jeff Schneider, Vince Nakayama
Hugging Face Trending Papers
Jun 18

See-and-Reach: Precise Vision-Language Navigation for UAVs within the Field of View

UAV Vision-Language Navigation (UAV-VLN) is typically formulated as a holistic search-and-reach problem, where long-range target discovery and final target approach are optimized and evaluated jointly. This formulation makes it difficult to assess a critical capability of aerial embodied agents, namely whether a UAV can accurately ground a visible target and translate vision-language evidence into precise 3D motion once the target enters its field of view.

arXiv AI
Jun 19

See-and-Reach: Precise Vision-Language Navigation for UAVs within the Field of View

arXiv:2606. 20045v1 Announce Type: cross Abstract: UAV Vision-Language Navigation (UAV-VLN) is typically formulated as a holistic search-and-reach problem, where long-range target discovery and final target approach are optimized and evaluated jointly.

By Fanfu Xue, En Yu, Yantian Shen, Zhikun Hu, Hongjun Wang, Yang Yang, Xindi Wang, Jiande Sun
arXiv AI
Sep 11

Decision Transformer for UAV-Mounted RIS-Assisted Dynamic D2D Communications

The paper investigates UAV‑mounted reconfigurable intelligent surface (RIS) assisted device‑to‑device (D2D) communication with stochastic link activation. It models UAV motion, attitude, time‑varying Rician angles, and angle‑dependent RIS reflection, and formulates a joint optimization of UAV trajectory, attitude, and RIS phases to maximize average sum rate under mobility, energy, and hardware constraints. The authors employ deep reinforcement learning and a Decision Transformer trained on expert trajectories from multiple scenarios, showing that zero‑shot transfer outperforms direct DRL transfer and that online fine‑tuning achieves competitive performance with fewer interactions.

By Yaxuan Liu
arXiv AI
Jul 9

Vision Language Action (VLA) Models for Unmanned Aerial Robotics and Bimanual Manipulation: A Review

arXiv:2607. 06706v1 Announce Type: cross Abstract: Vision Language Action (VLA) models unify visual perception, natural-language understanding, and action generation within a single foundation model, allowing a robot to follow instructions such as fold the towel or fly to the red building directly from camera images.

By Inkyu Sa, Chanoh Park, Hea-Min Lee, Donghee Noh, Ho Seok Ahn