arXiv AI

Demo: Vision-Language Model-Guided Online Calibration of an Electromagnetic Digital Twin

arXiv AI
23h ago

Seeing the Invisible: Physics-Guided Visual Prompting for Temperature- and Radiation-Aware VLA Navigation

The paper introduces Physics‑Guided Visual Prompting (PG‑VP), a plug‑and‑play module that overlays a virtual obstacle onto the input of a frozen Vision‑Language‑Action model to guide navigation around invisible hazards such as radiation or temperature spikes. PG‑VP performs a physics‑based risk assessment to determine the avoidance direction and dynamically renders the same virtual obstacle across frames, allowing the existing navigation policy to detour without retraining. Experiments on OmniNav with R2R‑CE and RxR‑CE datasets show that PG‑VP steers the policy toward low‑risk actions in 84.9% and 83.2% of cases, while real‑world tests on a robot demonstrate significant safety improvements against thermal and radiation sources.

By Hojoon Son, Fan Zhang
Hugging Face Trending Papers
Jul 22

NavVerse: Benchmarking Indoor-to-Outdoor Embodied Navigation in Continuous Robot Simulation

Robots deployed in delivery, campus, and emergency-response settings often need to navigate from buildings to streets within a single continuous episode. Existing benchmarks usually evaluate indoor and outdoor navigation separately, and many abstract away robot execution, leaving exit finding, boundary traversal, adaptation, and kinodynamic failures underexplored.

arXiv Computer Vision
Sep 2

Hydra: Marker-Free RGB-D Hand-Eye Calibration

Hydra introduces a marker‑free RGB‑D hand‑eye calibration method that leverages a novel ICP algorithm with a robust point‑to‑plane objective on a Lie algebra. Experiments on three serial manipulators and two RGB‑D cameras show that with only three random robot configurations the method achieves about 90% successful calibrations, 2–3× faster convergence to the global optimum, and 2 orders of magnitude faster convergence time (0.8 ± 0.4 s) compared to other marker‑free baselines. The approach delivers improved accuracy (5 mm in task space versus 7 mm for classical methods) while remaining marker‑free, and the authors provide an open‑source dataset, code, and ROS 2 integration.

By Martin Huber, Huanyu Tian, Christopher E. Mower, Lucas-Raphael M\"uller, S\'ebastien Ourselin, Christos Bergeles, Tom Vercauteren
arXiv Machine Learning
Sep 23

Bridging the Data Gap: Digital Twin as a New Paradigm for AI-based Radio Sensing

arXiv:2609.26214v1 Announce Type: new Abstract: We present a methodology that places a 3D digital twin (DT) of the environment as the main enabler behind the development of radio sensing at scale. Th...

By \'Eloi Sainte-Beuve (Orange Research), Guillaume Larue (Orange Research), Louis-Adrien Dufr\`ene (Orange Research), Quentin Lampin (Orange Research), Ali Al Khansa (Orange Research)
arXiv AI
23h ago

Sim-to-Real Transfer of Vision-Language Navigation in Continuous Environments Using an Ackermann-Steered Mobile Robot

The paper presents a vision‑language navigation system that transfers from simulation to a real Ackermann‑steered mobile robot without relying on navigation graphs or panoramic views. It uses a Cross‑Modal Attention architecture trained on simulated data and fine‑tuned with limited real‑world episodes, leveraging linear photometric adjustments and a camera‑LiDAR sensor suite. Evaluation with SPL and nDTW metrics shows robust, adaptable navigation in continuous environments.

By Chalindu Abeywansa, Sahan Gunasekara, Devindi De Silva, Seniru Dissanayake, Ranga Rodrigo, Peshala Jayasekara
arXiv Machine Learning
Aug 27

Minimalist Visual Inertial Odometry

The paper introduces a minimalist visual-inertial odometry system that uses only four downward-facing photodiodes with optical Gabor masks and an IMU to estimate motion for differential-drive robots. By jointly optimizing mask parameters and a Temporal Convolutional Network in a physically-grounded simulator, the model decodes speed from the photodiode signals and combines it with IMU angular speed to produce a continuous planar trajectory. Experiments on a prototype robot across indoor and outdoor terrains show that the system closely follows reference trajectories without real-world fine-tuning.

By Francesco Pasti, Jeremy Klotz, Nicola Bellotto, Shree K. Nayar
arXiv Computer Vision
5d ago

UniTrackPLA: Unified Panorama-Language-Action Model for Instruction-Guided Navigation and Dynamic Person Tracking

UniTrackPLA introduces a unified panorama-language-action model that simultaneously handles instruction‑guided navigation and dynamic person tracking for embodied robots. Its Panoramic‑Aware Encoding preserves azimuthal and temporal structure, allowing a shared vision‑language backbone to generate continuous waypoint chunks for both tasks. The model also employs World‑Action Consistency to predict future visual states and verify waypoint prefixes, enabling reliable action reuse and replanning when inconsistencies arise. A new OmniTrackNav‑Bench dataset and extensive real‑world experiments demonstrate significant performance gains over prior methods.

By Pengfei Qi, Haoran Lin, Sizhuang Chen, Kai Luo, Sirui Zhang, Xinqi Liu, Fei Cheng, Wenrui Chen, Liming Yin, Kailun Yang