arXiv:2511.16811v2 Announce Type: replace
Abstract: Building on the third-wave Extended Mind (EM) theory and radical enactivism, this article suggests an alternative to representation-based models of...
By Michael Carl, Takanori Mizowaki, Aishvarya Raj, Masaru Yamada, Devi Sri Bandaru, Yuxiang Wei, Xinyue Ren
arXiv:2609.10506v1 Announce Type: cross
Abstract: Action-conditioned latent world models predict future visual representations, enabling zero-shot goal-conditioned robot planning and control. However...
By Nisarga Nilavadi, Ralf R\"omer, Moritz Reuss, Michael Krawez, Tobias J\"ulg, Angela P. Schoellig, Rudolf Lioutikov, Wolfram Burgard
arXiv:2609.08636v1 Announce Type: cross
Abstract: Egocentric 4D interaction forecasting aims to anticipate both where future interactions will occur in 3D and how the human body will move to realize...
By Qiaohui Chu, Haoyu Zhang, Meng Liu, Haoxiang Shi, Dongmei Jiang, Liqiang Nie
arXiv:2609.09158v1 Announce Type: cross
Abstract: We study the problem of navigating cluttered indoor environments with a humanoid robot. Unlike conventional methods that model navigation as a 2D pat...
By Anqi Li, Yuxin Chen, Zhaobo Li, Zhuo Cao, Junli Ren, Masayoshi Tomizuka, Dhruv Shah
arXiv:2609.07738v1 Announce Type: cross
Abstract: LiDAR-based 3D Single Object Tracking (3D SOT) is critical for robotic perception and navigation and aims to localize dynamic objects across frames i...
By Zhaofeng Hu, Sifan Zhou, Jiahao Nie, Ziyu Zhao, Weizi Li, Ci-jyun Liang
MamMA is a pedestrian trajectory prediction algorithm that leverages LiDAR-generated occupancy maps and egocentric vision sensor data. It partitions the occupancy map into patches to extract obstacle features and incorporates pedestrian awareness states, which influence perception and speed. Using a Mamba-based model, MamMA predicts future trajectories and outperforms state‑of‑the‑art methods on multiple benchmark datasets.
By Juncen Long, Xiaofeng Jin, Gianluca Bardaro, Simone Mentasti, Matteo Matteucci
arXiv:2609.05519v1 Announce Type: cross
Abstract: We propose a unified strategy for fast goal inference in human-robot interaction. The core idea is to drive the human toward Critical Decision Points...
By Debasmita Ghose, Oz Gitelson, Michal Lewkowicz, Jake Brawer, Marynel Vazquez, Brian Scassellati
arXiv:2608.25757v4 Announce Type: replace-cross
Abstract: Large-scale vision--language--action (VLA) policies have advanced generalist robot control, yet most remain stimulus-to-action black boxes: a...
By Jin Lou, Zhiyuan Jing, Xupeng Wang, Andong Chen, Xingdong Zhu, Yuexuan Li, Yuan Xu, Zhijie Zhu, Yingwei Ji, Wenpeng Nie, Renxing Feng, Liangliang Chen, Ying Chu, Jingyi Li, Jinyan Liu, Zhiqi Song, Jingxuan Zhu, Jidong Zhang, Yufei Liu, Boyang Xing, Lei Jiang, Yan Cui, Hongming Li, Yuchen Zhu
arXiv:2609.09250v1 Announce Type: cross
Abstract: A verifier for robot policies reads a candidate behavior and returns a score for how well it did, used both to evaluate vision-language-action polici...
By Yang Wan, Xihang Yue, Zhirui Liu, Ziyuan Chu, Shuxun Wang, Yuhan Chen, Xiaonan Jiang, Xukun Zhu, Yubo Dong, Linchao Zhu
The paper introduces Distributed Dexterous Manipulation (DDM), a challenging control problem involving 64 soft delta robots arranged in an 8x8 grid. It presents a framework using spatially conditioned Multi-Agent Transformers (MATs) with adaptive layer norm, spatial contrastive embeddings, and a behavior cloning method fine‑tuned by Soft Actor Critic. Experiments demonstrate that MATs refine actions through stacked attention blocks, enabling long‑horizon planar manipulation in simulation and real‑world settings, while an action‑selection strategy reduces robot usage by about 65% and lowers wear‑and‑tear, achieving an average error of ~1.5 cm.
By Sarvesh Patil
arXiv:2605.03927v3 Announce Type: replace
Abstract: Vision-language models have demonstrated strong performance across robotic perception and instruction-following tasks. However, they still struggle...
By Xiaowen Sun, Matthias Kerzel, Mengdi Li, Xufeng Zhao, Paul Striker, Stefan Wermter
arXiv:2609.07741v1 Announce Type: new
Abstract: Persistent AI assistants are intended to extend human attention, memory, and coordination across changing digital and physical environments. To be trul...
By Jo\~ao Dias Ferreira
FALCON‑S is a modular, high‑fidelity simulator designed for fixed‑wing aerial robots operating near the ground. It models full 6DoF rigid‑body physics, semi‑empirical ground‑effect aerodynamics, actuator dynamics, sensor noise, and environmental disturbances, and supports both CPU and GPU backends via Torch and NVIDIA Warp for large‑scale reinforcement learning and optimal control. The framework offers a unified interface for various controllers, including RL and optical control algorithms, and allows cross‑validation with X‑Plane and JSBSim for engineering integration and visual fidelity.
By Matteo El Hariry, Pedro Lima, Andrej Orsula, Matthieu Geist, Miguel Olivares-Mendez
X2-NativeCursor is a lightweight observer that tracks text progress in incremental‑text streaming TTS by aligning native speech tokens to the original text before waveform decoding. It uses a normalization plan and a local matcher to estimate the current label position, converting these estimates into a non‑backtracking cursor. The method achieves a mean absolute error of 0.151 Chinese characters with 80‑ms lookahead, outperforming a waveform‑based baseline and reducing alignment real‑time factor from 0.3598 to 0.0180.
By Zehan Liu, Carl Chen, Rime Wen, Kaiqi Fu, Altman Lin, Shawn Qin, Lights Shi, Roy Gan, Hao Wang, Qian Wang
The Living Library is an end‑to‑end framework that converts fragmented digital archives into governed, conversational exhibit experiences. Developed at the Theodore Roosevelt Presidential Library, it digitizes a 300,000‑record collection, enriches it with OCR and metadata, and publishes it to a hybrid dense/semantic index. The system supports curator review via the Archivist App, powers a researcher interface, and runs Talk to TR—a museum exhibit where a digital human embodiment of Theodore Roosevelt answers visitors’ questions using Cross‑Era Analogical Grounding and dual‑path retrieval to keep responses grounded and responsive.
By Pengce Wang, Lucia Ronchi Darre, Matt Briney, Michaell Bakalars, Dan Rutkowski, Ursula Hardy, David Wolf, Laura Hoffman, Allen Kim, Shawn Wright, Juan Lavista Ferres
arXiv:2601.22475v2 Announce Type: replace
Abstract: Building a generalist robot policy requires continuously integrating new skills while preserving previously acquired behaviors. Directly optimizing...
By Qijun He, Yuxuan Li, Mingqi Yuan, Xiaoquan Sun, Wen-Tse Chen, Jeff Schneider, Jiayu Chen
arXiv:2609.06880v1 Announce Type: cross
Abstract: Reasoning over language instructions in embodied tasks such as robotics often requires understanding spatial relations from a speaker's situated pers...
By Mimo Shirasaka, Haochen Zhang, Yonatan Bisk
arXiv:2609.06195v1 Announce Type: new
Abstract: This paper proposes a transferable Map of Dynamics (MoD) framework that generalizes to unknown environments using only egocentric 3D LiDAR point clouds...
By Azusa Sawada, Allan Wang, Hideo Saito, Aaron Steinfeld
arXiv:2609.08084v1 Announce Type: cross
Abstract: Monocular depth estimation is a ubiquitous yet highly ill-posed computer vision task, with downstream applications in scene reconstruction, computati...
By Igor Pavlovic, Thiemo Wandel, Anton Obukhov, Luca Bartolomei, Andrey Davydov, Fabio Tosi, Matteo Poggi, Sabine S\"usstrunk, Dengxin Dai
arXiv:2509.06285v2 Announce Type: cross
Abstract: LiDAR point cloud registration is fundamental to robotic perception and navigation. In geometrically degenerate environments (e.g., corridors), regis...
By Xiangcheng Hu, Xieyuanli Chen, Mingkai Jia, Jin Wu, Ping Tan, Steven L. Waslander