Robotics and embodied AI

Manipulation, locomotion, sim-to-real transfer and autonomous driving: learning systems that have to survive physics.

3,943 stories · RSS feed

arXiv Machine Learning
Sep 10

MamMA: A Mamba-Based Pedestrian Trajectory Prediction Algorithm Considering Occupancy Map and Pedestrian Awareness States

MamMA is a pedestrian trajectory prediction algorithm that leverages LiDAR-generated occupancy maps and egocentric vision sensor data. It partitions the occupancy map into patches to extract obstacle features and incorporates pedestrian awareness states, which influence perception and speed. Using a Mamba-based model, MamMA predicts future trajectories and outperforms state‑of‑the‑art methods on multiple benchmark datasets.

By Juncen Long, Xiaofeng Jin, Gianluca Bardaro, Simone Mentasti, Matteo Matteucci
arXiv Machine Learning
Sep 10

LM-X: Explainable Vision--Language--Action Modeling via Progress, Event, and Uncertainty Prediction

arXiv:2608.25757v4 Announce Type: replace-cross Abstract: Large-scale vision--language--action (VLA) policies have advanced generalist robot control, yet most remain stimulus-to-action black boxes: a...

By Jin Lou, Zhiyuan Jing, Xupeng Wang, Andong Chen, Xingdong Zhu, Yuexuan Li, Yuan Xu, Zhijie Zhu, Yingwei Ji, Wenpeng Nie, Renxing Feng, Liangliang Chen, Ying Chu, Jingyi Li, Jinyan Liu, Zhiqi Song, Jingxuan Zhu, Jidong Zhang, Yufei Liu, Boyang Xing, Lei Jiang, Yan Cui, Hongming Li, Yuchen Zhu
arXiv Machine Learning
Sep 10

Distributed Dexterous Manipulation with Spatially Conditioned Multi-Agent Transformers

The paper introduces Distributed Dexterous Manipulation (DDM), a challenging control problem involving 64 soft delta robots arranged in an 8x8 grid. It presents a framework using spatially conditioned Multi-Agent Transformers (MATs) with adaptive layer norm, spatial contrastive embeddings, and a behavior cloning method fine‑tuned by Soft Actor Critic. Experiments demonstrate that MATs refine actions through stacked attention blocks, enabling long‑horizon planar manipulation in simulation and real‑world settings, while an action‑selection strategy reduces robot usage by about 65% and lowers wear‑and‑tear, achieving an average error of ~1.5 cm.

By Sarvesh Patil
arXiv Machine Learning
Sep 10

FALCON-S: Fixed-wing ground-effect Aerodynamics Simulator and Flight Control Learning Suite

FALCON‑S is a modular, high‑fidelity simulator designed for fixed‑wing aerial robots operating near the ground. It models full 6DoF rigid‑body physics, semi‑empirical ground‑effect aerodynamics, actuator dynamics, sensor noise, and environmental disturbances, and supports both CPU and GPU backends via Torch and NVIDIA Warp for large‑scale reinforcement learning and optimal control. The framework offers a unified interface for various controllers, including RL and optical control algorithms, and allows cross‑validation with X‑Plane and JSBSim for engineering integration and visual fidelity.

By Matteo El Hariry, Pedro Lima, Andrej Orsula, Matthieu Geist, Miguel Olivares-Mendez
arXiv Computation and Language
Sep 10

X2-NativeCursor: Native-Token Text Progress Tracking for Incremental-Text Streaming Codec TTS

X2-NativeCursor is a lightweight observer that tracks text progress in incremental‑text streaming TTS by aligning native speech tokens to the original text before waveform decoding. It uses a normalization plan and a local matcher to estimate the current label position, converting these estimates into a non‑backtracking cursor. The method achieves a mean absolute error of 0.151 Chinese characters with 80‑ms lookahead, outperforming a waveform‑based baseline and reducing alignment real‑time factor from 0.3598 to 0.0180.

By Zehan Liu, Carl Chen, Rime Wen, Kaiqi Fu, Altman Lin, Shawn Qin, Lights Shi, Roy Gan, Hao Wang, Qian Wang
arXiv Computer Vision
Sep 10

The Living Library: Transforming Archival Collections into Conversational Knowledge Systems -- Lessons from the Theodore Roosevelt Presidential Library

The Living Library is an end‑to‑end framework that converts fragmented digital archives into governed, conversational exhibit experiences. Developed at the Theodore Roosevelt Presidential Library, it digitizes a 300,000‑record collection, enriches it with OCR and metadata, and publishes it to a hybrid dense/semantic index. The system supports curator review via the Archivist App, powers a researcher interface, and runs Talk to TR—a museum exhibit where a digital human embodiment of Theodore Roosevelt answers visitors’ questions using Cross‑Era Analogical Grounding and dual‑path retrieval to keep responses grounded and responsive.

By Pengce Wang, Lucia Ronchi Darre, Matt Briney, Michaell Bakalars, Dan Rutkowski, Ursula Hardy, David Wolf, Laura Hoffman, Allen Kim, Shawn Wright, Juan Lavista Ferres
arXiv Machine Learning
Sep 10

Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation

arXiv:2609.08084v1 Announce Type: cross Abstract: Monocular depth estimation is a ubiquitous yet highly ill-posed computer vision task, with downstream applications in scene reconstruction, computati...

By Igor Pavlovic, Thiemo Wandel, Anton Obukhov, Luca Bartolomei, Andrey Davydov, Fabio Tosi, Matteo Poggi, Sabine S\"usstrunk, Dengxin Dai