Robotics and embodied AI

Manipulation, locomotion, sim-to-real transfer and autonomous driving: learning systems that have to survive physics.

3,943 stories · RSS feed

arXiv Machine Learning
Sep 16

EMODY Flow: Emotion-Aware Audio-Driven Full-Body Motion Generation

EMODY Flow is a lightweight flow‑matching framework that generates synchronized full‑body motion and facial expressions conditioned on speech and emotion. It attaches to a frozen Qwen‑3 Omni model, reusing its audio codecs to drive two parallel DiT generators for SMPL‑X body pose and FLAME facial expressions. An auxiliary emotion classifier at training time restores emotion sensitivity, enabling EMODY Flow to achieve state‑of‑the‑art gesture quality on BEAT2 and zero‑shot facial animation on TFHP, with significant improvements in FGD, Beat Correlation, and Diversity metrics.

By Harsh Kumar Agarwal, Xavier Alameda-Pineda, Olivier Perrotin
arXiv Computer Vision
Sep 16

NeuroSymbEAD: A Large Scale Neuro-Symbolic Caption Dataset for Omni-Directional Embodied Autonomous Driving

NeuroSymbEAD is a large‑scale neuro‑symbolic caption dataset that builds an ego‑centric knowledge graph of static and dynamic objects on the KITTI‑360 dataset, annotating classes, categories, heading directions, orientations, and distances from the ego‑vehicle. The dataset generates multilevel textual captions that serve as a lightweight representation of an ego‑centric scene map, enabling outdoor scene‑map reconstruction, visual recognition, and object grounding. Baselines for driving common sense and traffic/scene understanding are established, and the dataset is benchmarked using pre‑trained grounding and learned auto‑regressive captioning networks to support vision‑language and foundation models for traffic‑scene explanation, 3D reasoning, and interpretable autonomous‑driving perception.

By Muhammad Ahmed Ullah Khan, Mohammed Elamine, Sheikh Talha Uddin, Didier Stricker, Sk Aziz Ali, Muhammad Zeshan Afzal
arXiv Computation and Language
Sep 16

surprisal is Not a Theory

The article argues that Surprisal Theory, often presented as a computational-level explanation, is not a theory in its own right. It contends that using large language model (LLM) surprisals without considering the underlying representational and algorithmic choices obscures the theory’s commitments. The authors demonstrate through three analyses that algorithm and architecture significantly influence language model probabilities, urging researchers to reassess treating LLM surprisals as interchangeable.

By Andr\'es Bux\'o-Lugo, Aniello De Santo, Morgan Grobol, Ryan J. Hubbard, Cassandra L. Jacobs
arXiv Computer Vision
Sep 16

Exploring 2D backbone effects for indoor semantic occupancy prediction

The paper investigates how different 2D image backbones affect indoor semantic occupancy prediction in RGB‑D pipelines. By keeping the projection, depth branch, and occupancy head constant and swapping only the image encoder, the authors find that stronger backbones such as DINOv2 and BLIP2 significantly raise mIoU scores compared to the default ResNet‑50. These results show that the choice of image backbone is a major determinant of 3D occupancy accuracy, outweighing many specialized 3D modules.

By Shizhang Fanga, Wanling Yea, Qi Zheng
arXiv Computer Vision
Sep 16

Racing in Volume with Flow Ensembles

The paper introduces FastFlowGS, a streaming 4D Gaussian Splatting method that reconstructs fast-moving subjects from a sparse set of external cameras, and Monaco4D, a photorealistic Unreal Engine 5 benchmark featuring Formula 1 sequences with dense ground truth. FastFlowGS combines sparse matches, semi-dense tracks, and dense optical flow using a Kalman-style temporal update, achieving significant performance gains over existing baselines on both CMU-Panoptic and Monaco4D datasets. The benchmark provides varied illumination and viewpoints from trackside, onboard, and drone cameras, enabling evaluation of high-speed outdoor reconstruction.

By Saswat Subhajyoti Mallick, Riu Cherdchusakulchai, Marc Ruiz Olle, Albert Mosella-Montoro, Jose Ribeiro-Gomes, Francisco Vicente Carrasco, Fernando De la Torre
arXiv Computer Vision
Sep 16

Reasoning with Image Generation

The paper introduces ReImaGin, a method that uses image generation models as a flexible visual reasoning tool for multimodal large language models. Unlike traditional fixed-function vision tools, ReImaGin accepts natural language commands and can perform open-ended visual operations such as removing occlusions or creating floorplans from multiple views. Experiments on six diverse visual reasoning tasks show that ReImaGin outperforms both text-only reasoning and specialist vision-tool baselines, achieving up to a 25% improvement.

By Nishad Singhi, Hector Garcia Rodriguez, Aditya Arora, Marcus Rohrbach, Anna Rohrbach
arXiv Computer Vision
Sep 16

MEgoVista: Multi-view Ego-aware Motion Estimation for Metric 4D Hands and Head in the Wild

MEgoVista is an offline pipeline that converts a single unprepared egocentric video into metric two‑hand and head motion within a gravity‑aligned world frame. It uniquely reconstructs motion in environments beyond studio volumes, uses calibrated stereo for absolute scale, and evaluates its outputs against independent optical capture to audit accuracy. The system thus expands the settings where high‑fidelity hand‑motion labels can be generated from natural, head‑worn recordings.

By Jiangong Xiao (Northwestern Polytechnical University), Zhihao Zhang (Xi'an Jiaotong University), Yifei Dong (Maniformer), Chao Ma (Maniformer), Zhouyi Jin (Maniformer), Zhiwen Hou (Maniformer), Li Liu (Maniformer), Weihuang Chen (Xi'an Jiaotong University), Hongbin Sun (Xi'an Jiaotong University), Maoqing Yao (Maniformer)
arXiv Computer Vision
Sep 16

GeoLAM: Learning Geometry-Grounded Latent Actions from Unlabeled Human Videos

GeoLAM is a framework that learns geometry‑grounded latent actions from unlabeled human videos. It uses future‑frame reconstruction with a frozen geometric feature hierarchy and motion supervision from a 4D geometry teacher to capture 3D displacement, image‑plane motion, and surface‑orientation changes. After pretraining, the representation serves as transition targets for a world‑action model trained on robot demonstrations, enabling denoised latent actions and executable action chunks without requiring hand‑pose annotations or future‑video generation during deployment.

By Yifan Xie, Hekun Tian, Jinkun Liu, YuAn Wang, Qiao Sun, Wenbo Ding
arXiv Machine Learning
Sep 16

The Neverwhere Visual Parkour Benchmark Suite

The paper introduces the Neverwhere Visual Parkour Benchmark Suite, a collection of over sixty hyper‑photo‑realistic 3D Gaussian Splatting reconstructions of urban indoor and outdoor scenes designed to evaluate visual locomotion controllers in closed‑loop, continuous testing setups. It aims to bridge the gap between training and real‑world evaluation by providing reproducible environments and policy checkpoints trained across multiple scenes, while highlighting the risks of relying solely on 3D Gaussian‑generated data. The authors offer code and data on their project page for easy integration into robotic evaluation pipelines.

By Ziyu Chen, Henghui Bao, Haoran Chang, Alan Yu, Ran Choi, Kai McClennen, Gio Huh, Kevin Yang, Ri-Zhao Qiu, Yajvan Ravan, John J. Leonard, Xiaolong Wang, Phillip Isola, Ge Yang, Yue Wang
arXiv Machine Learning
Sep 16

Intrinsic Robot Rewarding: Reusing VLA Representations for Autonomous Evaluation and Policy Improvement

Intrinsic Robot Rewarding (IRR) leverages existing vision‑language‑action (VLA) systems to evaluate a robot’s own outcomes and provide feedback for policy improvement. By using successful demonstration endpoints as task‑specific references and the policy’s frozen visual encoder as the feature space, IRR adds a reference bank and scoring operation to the current pipeline without requiring a separate evaluator or additional perception backbone. The approach aims to reduce integration effort, reward computation cost, and human outcome scoring while enabling learning from the data already available in industrial robot systems.

By Tobias Schaffer, Mohab Elkhayat, Daniela Nicklas, Mustafa Almohamad, Elham Al-Fuqara
arXiv Machine Learning
Sep 16

Digital Persuasion: Understanding the Impact of Online Influencers on Public Opinion

The paper proposes a Friedkin‑Johnsen based framework to identify influential users in online social networks and assess how they shape community opinion. By manipulating initial opinions in experiments, the authors show that top influencers can significantly shift overall community sentiment, and their influence extends beyond direct neighbors to second‑degree contacts. The framework is validated on a tweet dataset from the U.S. presidential election, illustrating the power of digital influencers to alter public opinion.

By Omran Berjawi, Rida Khatoun, Giuseppe Fenza
arXiv AI
Sep 16

SafeFlow: Real-Time Text-Driven Humanoid Whole-Body Control via Physics-Guided Rectified Flow and Selective Safety Gating

SafeFlow is a real‑time, text‑driven humanoid control framework that blends physics‑guided motion generation with a three‑stage safety gate. It uses Physics‑Guided Rectified Flow Matching in a VAE latent space to produce physically executable trajectories, accelerates sampling with Reflow, and filters unsafe outputs via semantic OOD detection, directional sensitivity checks, and hard kinematic constraints before handing them to a motion‑tracking controller. Experiments on the Unitree G1 show that SafeFlow achieves higher success rates, better physical compliance, and faster inference than diffusion‑ and retargeting‑based baselines while maintaining motion diversity.

By Hanbyel Cho, Sang-Hun Kim, Jeonguk Kang, Donghan Koo
arXiv AI
Sep 16

CTAN: Cycle-Temporal Attention Network for Embodied Audio-Visual Navigation

The paper introduces CTAN, a Cycle-Temporal Attention Network for audio‑visual embodied navigation. It proposes an Audio‑Visual Reconstruction Cross‑Attention module that uses bidirectional cycle‑consistency to strengthen spatial semantics across visual and acoustic modalities, and a Temporal Cross‑Modal Memory to fuse real‑time multimodal features with historical context. Experiments on Replica and Matterport3D show that CTAN outperforms prior methods in success rate, SPL, and scene navigation accuracy.

By Teng Liu, Yinfeng Yu