Robotics and embodied AI

Manipulation, locomotion, sim-to-real transfer and autonomous driving: learning systems that have to survive physics.

3,943 stories · RSS feed

arXiv Computer Vision
Sep 16

Reasoning with Image Generation

The paper introduces ReImaGin, a method that uses image generation models as a flexible visual reasoning tool for multimodal large language models. Unlike traditional fixed-function vision tools, ReImaGin accepts natural language commands and can perform open-ended visual operations such as removing occlusions or creating floorplans from multiple views. Experiments on six diverse visual reasoning tasks show that ReImaGin outperforms both text-only reasoning and specialist vision-tool baselines, achieving up to a 25% improvement.

By Nishad Singhi, Hector Garcia Rodriguez, Aditya Arora, Marcus Rohrbach, Anna Rohrbach
arXiv Computer Vision
Sep 16

NeuroSymbEAD: A Large Scale Neuro-Symbolic Caption Dataset for Omni-Directional Embodied Autonomous Driving

NeuroSymbEAD is a large‑scale neuro‑symbolic caption dataset that builds an ego‑centric knowledge graph of static and dynamic objects on the KITTI‑360 dataset, annotating classes, categories, heading directions, orientations, and distances from the ego‑vehicle. The dataset generates multilevel textual captions that serve as a lightweight representation of an ego‑centric scene map, enabling outdoor scene‑map reconstruction, visual recognition, and object grounding. Baselines for driving common sense and traffic/scene understanding are established, and the dataset is benchmarked using pre‑trained grounding and learned auto‑regressive captioning networks to support vision‑language and foundation models for traffic‑scene explanation, 3D reasoning, and interpretable autonomous‑driving perception.

By Muhammad Ahmed Ullah Khan, Mohammed Elamine, Sheikh Talha Uddin, Didier Stricker, Sk Aziz Ali, Muhammad Zeshan Afzal
arXiv Computer Vision
Sep 16

GeoLAM: Learning Geometry-Grounded Latent Actions from Unlabeled Human Videos

GeoLAM is a framework that learns geometry‑grounded latent actions from unlabeled human videos. It uses future‑frame reconstruction with a frozen geometric feature hierarchy and motion supervision from a 4D geometry teacher to capture 3D displacement, image‑plane motion, and surface‑orientation changes. After pretraining, the representation serves as transition targets for a world‑action model trained on robot demonstrations, enabling denoised latent actions and executable action chunks without requiring hand‑pose annotations or future‑video generation during deployment.

By Yifan Xie, Hekun Tian, Jinkun Liu, YuAn Wang, Qiao Sun, Wenbo Ding
arXiv Machine Learning
Sep 16

Intrinsic Robot Rewarding: Reusing VLA Representations for Autonomous Evaluation and Policy Improvement

Intrinsic Robot Rewarding (IRR) leverages existing vision‑language‑action (VLA) systems to evaluate a robot’s own outcomes and provide feedback for policy improvement. By using successful demonstration endpoints as task‑specific references and the policy’s frozen visual encoder as the feature space, IRR adds a reference bank and scoring operation to the current pipeline without requiring a separate evaluator or additional perception backbone. The approach aims to reduce integration effort, reward computation cost, and human outcome scoring while enabling learning from the data already available in industrial robot systems.

By Tobias Schaffer, Mohab Elkhayat, Daniela Nicklas, Mustafa Almohamad, Elham Al-Fuqara
arXiv Machine Learning
Sep 16

The Neverwhere Visual Parkour Benchmark Suite

The paper introduces the Neverwhere Visual Parkour Benchmark Suite, a collection of over sixty hyper‑photo‑realistic 3D Gaussian Splatting reconstructions of urban indoor and outdoor scenes designed to evaluate visual locomotion controllers in closed‑loop, continuous testing setups. It aims to bridge the gap between training and real‑world evaluation by providing reproducible environments and policy checkpoints trained across multiple scenes, while highlighting the risks of relying solely on 3D Gaussian‑generated data. The authors offer code and data on their project page for easy integration into robotic evaluation pipelines.

By Ziyu Chen, Henghui Bao, Haoran Chang, Alan Yu, Ran Choi, Kai McClennen, Gio Huh, Kevin Yang, Ri-Zhao Qiu, Yajvan Ravan, John J. Leonard, Xiaolong Wang, Phillip Isola, Ge Yang, Yue Wang
arXiv Machine Learning
Sep 16

Digital Persuasion: Understanding the Impact of Online Influencers on Public Opinion

The paper proposes a Friedkin‑Johnsen based framework to identify influential users in online social networks and assess how they shape community opinion. By manipulating initial opinions in experiments, the authors show that top influencers can significantly shift overall community sentiment, and their influence extends beyond direct neighbors to second‑degree contacts. The framework is validated on a tweet dataset from the U.S. presidential election, illustrating the power of digital influencers to alter public opinion.

By Omran Berjawi, Rida Khatoun, Giuseppe Fenza