Robotics and embodied AI

Manipulation, locomotion, sim-to-real transfer and autonomous driving: learning systems that have to survive physics.

3,855 stories · RSS feed

arXiv AI
4d ago

FineART: Fine-Grained Annotated Robotic Trajectory Dataset and Vision-Language-Action Model for Bimanual Manipulation

arXiv:2609.36416v2 Announce Type: replace-cross Abstract: Robots operating in real-world environments must often execute complex, multi-step bimanual tasks over long horizons rather than single, isol...

By Jade Choghari, Pepijn Kooijmans, Mansi Agarwal, Yusuf Umut Ciftci, Aseem Doriwala, Catherine Weaver, Mouli Sivapurapu, Kai Yang, Thomas Wolf, Jackson Lee, Pragna Mannam
arXiv Machine Learning
4d ago

STARS: From Spatiotemporal Dynamics to Social Representations in Human-Robot Interaction

arXiv:2609.40245v2 Announce Type: cross Abstract: Robot navigation in dynamic, human-centered environments requires socially-compliant decisions grounded in robust scene understanding. Recent Vision-...

By Nathan Tsoi, Michael J. Munje, Tejas Oberoi, Rishab Maheshwari, Pengen Zheng, Tanush Chauhan, Peter Stone, Joydeep Biswas
arXiv AI
4d ago

Screw Attention: Rigid-Body Algebra Inside a Transformer

Screw Attention introduces a transformer layer that treats the relation between two bodies as a spatial transform rather than a graph edge, enabling each token to represent a body with a pose and its relative pose or joint screw. This design ensures equivariance to independent frame changes and allows a single layer to capture rigid‑body velocity recursion. Experiments on simulated manipulation tasks show that Screw Attention matches or outperforms other network architectures, achieving high success rates on LIBERO‑Spatial with far fewer parameters and maintaining performance under frame convention changes and pose noise.

By Aly Magassouba
arXiv AI
5d ago

Diffusion-Augmented Markov Decision Processes for Maximum Entropy Reinforcement Learning

The paper introduces Diffusion-Augmented Markov Decision Processes (DA‑MDPs), a framework that extends Maximum Entropy Reinforcement Learning to diffusion-based policies. DA‑MDPs treat each reverse‑diffusion step as an RL decision, deriving a tractable reverse‑KL bound that decomposes across denoising transitions and yields diffusion‑augmented soft rewards, value functions, and policy objectives. The authors implement this framework with PPO, REPPO, and a maximum‑entropy WPO variant, showing improved continuous‑control performance, higher success rates on manipulation tasks, and memory‑efficient training with action chunking.

By Sebastian Sanokowski, Kaustubh Patil, Majid Khadiv
arXiv AI
5d ago

FoCLIP: A Feature-Space Misalignment Framework for CLIP-Based Image Manipulation and Detection

FoCLIP is a framework that creates adversarial examples to manipulate CLIP-based image quality metrics by reducing the alignment between image and text features. It uses stochastic gradient descent to combine feature alignment, score distribution balancing, and pixel‑guard regularization, enabling high CLIPscore predictions while maintaining visual fidelity. Experiments on artistic prompts and ImageNet show significant CLIPscore gains, and the authors also propose a color‑channel sensitivity detection method that achieves 91% accuracy.

By Yulin Chen, Zeyuan Wang, Tianyuan Yu, Yingmei Wei, Liang Bai
arXiv AI
5d ago

Efficient Multi-Modal Planning with Reward-Guided Preference Optimization for Autonomous Driving

Efficient Multi-Modal Planning with Reward-Guided Preference Optimization for Autonomous Driving proposes EMPlan, a hybrid trajectory planning method that combines sparse anchors with an offset refinement module for low-latency, high-accuracy predictions. The approach uses a two-stage training paradigm—pretraining followed by reward-guided fine-tuning—to improve safety without extra inference cost, leveraging rule-based reward signals and unpaired preference supervision. EMPlan is evaluated on the non-reactive NAVSIM benchmark, achieving a favorable balance between planning accuracy and efficiency under real-time constraints.

By Chenglin Chen, Lujia Wang, Xinhu Zheng, Jun Ma, Haoang Li
arXiv AI
5d ago

DiffWAM: A Fast and Efficient Navigation World Action Model

DiffWAM is a geometry‑conditioned navigation world‑action model that transforms predictive features from a frozen video foundation model into continuous camera trajectories, eliminating the need for future‑video synthesis and multi‑frame reconstruction during deployment. Its Grid‑Motion module preserves spatial‑temporal motion associations, while Latent2Pose grounds them with first‑frame geometry to recover metrically meaningful 3D motion. The system, complemented by FastDreamer for asynchronous trajectory handoff, achieves a trajectory RMSE of 0.3492 m and a 74.40 % endpoint success rate on the DiffWAM‑1000 benchmark, with real‑world tests showing complex UAV behaviors and an onboard implementation reaching 1.08 s latency on NVIDIA Jetson AGX Thor.

By Mo Zhu, Yuze Wu, Xijie Huang, Xiao Cui, Fei Gao, Xin Zhou