Latent world models that integrate a flow in a frozen self supervised latent space train stably and cheaply, yet silently lose the property manipulation depends on most: motion. The pretrained flow ne...
Claims about AI safety reach audiences well beyond the AI community, yet many rely on opaque evidence or static assessments, when supporting evidence is accessible at all. We present the Systemic Risk...
Microsoft Research reports that offloading AI inference from robots to external hardware can enhance task success, increase efficiency, and enable more advanced physical AI workloads. The study suggests that moving inference beyond the robot’s onboard processors allows the hardware to keep pace with growing AI capabilities. This approach demonstrates a practical way to improve robotic performance in real-world settings.
By Ganesh Ananthanarayanan, Matthew Balkwill, Xenofon Foukas, Sanjeev Mehrotra, Bozidar Radunovic, Connor Settle, Ankit Verma, David White, Shawn Cicoria, Mark Martin, Rachel Johnson, Mayur Patel
We present VGM-VS, a visual servoing method built on a pretrained feed-forward visual geometry model. Given the current view and a reference image captured at the target configuration, we estimate the...
Robotic assembly in high-mixture settings requires adaptable systems that can handle diverse parts, yet current approaches typically rely on policies specialized to each insertion task. Although this...
With the development of applications like autonomous driving, object detection has gained significant attention, while also highlighting critical vulnerabilities like backdoor attacks that severely co...
Physical intelligence requires more than predicting how the world may evolve: predictions must remain actionable as the world continues to change. We introduce InternW0, the first instantiation of the...
Long-term autonomy in human-populated environments requires anticipating whether and how people will move at times a robot has not yet observed. Existing representations of pedestrian motion face a tr...
arXiv:2609.26567v1 Announce Type: cross
Abstract: Robots that store past experiences must select which one to reuse in a new scene. Most systems select by visual similarity, and most evaluations repo...
By Eshika Pathak, Leela Krishna
arXiv:2609.24124v1 Announce Type: cross
Abstract: Active perception and manipulation are crucial for robots to interact with complex scenes. Existing benchmarks struggle to evaluate how robots effect...
By Yibo Li, Enshen Zhou, Rui Chen, Yanjun Ding, Mengzhen Liu, Yi Han, Jiabo Zhan, Lipeng Wang, Shanghang Zhang, Lu Sheng
arXiv:2609.24815v2 Announce Type: cross
Abstract: Scalable simulation is essential for robot data generation, policy training, evaluation, and safe iteration, yet real-world interaction is costly and...
By Wenkang Qin, Yukun Zhou, Noah Shen, Jisong Cai, Dongxiao Mao, Baicheng Li, Yue Zhang, Wei Sui
ForeDrive introduces a planning-relevant latent world model that is asymmetrically coupled to a Diffusion Transformer planner. The model learns multi‑horizon latent futures with a JEPA‑style world model, while planning gradients update the shared encoder and stop‑gradient routing trains the predictor with forecasting losses only. Gated visual fusion, future‑status injection, and Trajectory‑Adaptive Bias are used to guide trajectory generation without overriding current observations, achieving high performance on NAVSIM benchmarks using only front‑view images and pure imitation learning.
By Sinuo Wang, Zichong Gu, Yuhan Huang, Wenxin Wen, Xun Yang, Yiqing Zhang, Xingyu Zhang, Ningyu Che, Jie Ling, Qiankun Yu, Wei Liu, Jing Xu, Xinggang Wang
The paper introduces ORDER, a fictitious-world benchmark designed to evaluate domain-adaptive embodied AI. ORDER consists of a synthetic 342,069-token corpus defining a self-consistent physics, a 500-question knowledge test (ORDER‑BENCH), and a compositional spatial task (ORDER‑SPATIAL) that requires ordering objects for safe manipulation. The benchmark demonstrates that models like GPT‑4.1 perform poorly without adaptation, while small models improve significantly after continual pre‑training, and that performance on ORDER‑SPATIAL better predicts real plan quality than knowledge-test accuracy.
By Sai Krishna Reddy Sathi, Anuj Tiwari
Tactile-JEPA is a self‑supervised pre‑training method for distributed tactile sensors that leverages the sensors’ spatial topology to learn topology‑aware representations. It predicts embeddings of masked sensing elements using a sensor connectivity graph and dual‑scale masking to capture both local contact details and the global tactile surface state. Evaluated on three diverse datasets, it improves force estimation by 6.3 % and in‑hand orientation error by 20.8 % over previous state‑of‑the‑art methods, and yields consistent gains in downstream tasks such as policy learning.
By Elizaveta Kovtun, Matvey Konovalov, Andrey Sakhovskiy, Semen Budennyy
The paper introduces an LLM-based Conversational AI Knowledge Assistant for the Raspberry‑Pi‑powered 13‑Axis MyBuddy humanoid robot. It combines large language model-driven language understanding, real‑time speech recognition, internet‑based knowledge retrieval (e.g., Wikipedia, arXiv), flexible dialogue management, and natural speech synthesis to support intelligent, multi‑turn conversations and emotional‑support interactions. This system aims to overcome the limitations of traditional rule‑based dialogue systems in humanoid robots.
By Hanxiao Chen
The paper introduces NavSecure, a vision‑based navigation framework that combines world models with safety‑focused decision making for autonomous vehicles operating over 5G networks. By predicting potential threats and generating safer routes, NavSecure reduces reliance on extensive real‑world trial‑and‑error learning and enables continuous autonomous adaptation. Experiments in simulation‑to‑real driving scenarios demonstrate superior performance in collision prevention and risk reduction compared to other end‑to‑end methods.
By Hong Ding, Ziming Wang, Yi Ding, Hongjie Lin, SuYang Xi, Chia Chao Kang
SAM‑V is a geometry‑aware extension of the Segment Anything Model (SAM) that integrates 3D priors from a feed‑forward geometry model (VGGT) into 2D segmentation. It uses a prompt‑fusion mechanism to combine sparse SAM prompts with view‑specific camera tokens and local VGGT features, enabling a mask decoder that attends to both dense 2D and 3D cues. The resulting end‑to‑end system produces consistent multi‑view instance segmentation in a single forward pass, achieving significant gains on the IGGT 3D tracking benchmark without offline mask matching or explicit 3D reconstruction.
By Jiangshan Gong, Yuqun Wu, Qiqian Fu, Yao Xiao, Chuhang Zou, Shenlong Wang, Derek Hoiem
MAVP (Map-Aware Visuomotor Policies) is a framework that enhances mobile manipulation by predicting explicit base-pose targets and tracking them with localisation feedback. It reconstructs a static map from teleoperated demonstrations, expressing base trajectories in a shared map frame to provide consistent spatial supervision. During execution, the policy uses RGB observations, joint states, and the robot’s current map-frame base pose to jointly predict target base poses, arm actions, and gripper actions, while a low‑level controller corrects deviations using feedforward motion and pose error feedback. Pose‑noise augmentation during training further improves robustness, and MAVP outperforms unanchored velocity control across six real‑world tasks and three policy families.
By Jinhe Tang, Ruixiao Dai, Weiming Zhi
The paper presents a foundation-guided auto‑annotation pipeline that improves standard autonomous driving object detectors in adverse weather. By benchmarking YOLOv8, Co‑DETR, and SAM3 on a custom dataset of 25 operational scenarios, the authors find SAM3 to be the most robust and use it offline to generate pseudo‑labels. Fine‑tuning YOLOv8 on these labels boosts overall mAP by 16.04% and yields significant gains in specific conditions such as Residential Direct Sunlight (32.73%) and Highway Fog (28.65%).
By Sepideh Gohari, Goodarz Mehr, Azim Eskandarian
Fysiverse-3D-Vision is a unified vision‑language‑geometry framework that reconstructs executable 3D scenes from a single image. It separates spatial layout reasoning from asset synthesis, using a shared representation where spatial reasoning and geometric reconstruction reinforce each other. The model employs a Transformer that integrates textual supervision, semantic visual cues, and geometric representations, and includes an object‑conditioned layout module to predict object translation, rotation, and scale while maintaining physical consistency through collision‑aware optimization.
By Dingkang Yang, Yizhou Liu, Wendong Cheng, Zizhi Chen, Shunli Wang, Yang Liu, Hongsheng Li, Lihua Zhang