arXiv:2609.09396v1 Announce Type: new
Abstract: As Vision-Language Models (VLMs) advance toward physical deployment, the focus has remained on action-oriented Embodied AI evaluated on subject-centric...
By Zaid Pervaiz Bhat, Nimra Nayyar, Arihant Jain, Lap Fung Chan, John Suchanek, Yu Wang, Varun Praveen, Tomasz Kornuta, Vidya Nariyambut Murali
arXiv:2609.09210v1 Announce Type: cross
Abstract: Teleoperated demonstrations are often multimodal even when the underlying dynamics are nearly deterministic given the executed action. We argue that...
By Jinting Hang, Zhenhui Cai
arXiv:2605.01234v2 Announce Type: replace
Abstract: We present TT4D, a large-scale, high-fidelity table tennis dataset. It provides $140+$ hours of reconstructed singles and doubles gameplay from mon...
By Nima Rahmanian, Daniel Kienzle, Thomas Gossard, Dvij Kalaria, Rainer Lienhart, Shankar Sastry
arXiv:2511.16811v2 Announce Type: replace
Abstract: Building on the third-wave Extended Mind (EM) theory and radical enactivism, this article suggests an alternative to representation-based models of...
By Michael Carl, Takanori Mizowaki, Aishvarya Raj, Masaru Yamada, Devi Sri Bandaru, Yuxiang Wei, Xinyue Ren
arXiv:2608.25757v4 Announce Type: replace-cross
Abstract: Large-scale vision--language--action (VLA) policies have advanced generalist robot control, yet most remain stimulus-to-action black boxes: a...
By Jin Lou, Zhiyuan Jing, Xupeng Wang, Andong Chen, Xingdong Zhu, Yuexuan Li, Yuan Xu, Zhijie Zhu, Yingwei Ji, Wenpeng Nie, Renxing Feng, Liangliang Chen, Ying Chu, Jingyi Li, Jinyan Liu, Zhiqi Song, Jingxuan Zhu, Jidong Zhang, Yufei Liu, Boyang Xing, Lei Jiang, Yan Cui, Hongming Li, Yuchen Zhu
arXiv:2609.08084v1 Announce Type: cross
Abstract: Monocular depth estimation is a ubiquitous yet highly ill-posed computer vision task, with downstream applications in scene reconstruction, computati...
By Igor Pavlovic, Thiemo Wandel, Anton Obukhov, Luca Bartolomei, Andrey Davydov, Fabio Tosi, Matteo Poggi, Sabine S\"usstrunk, Dengxin Dai
SpatialBlock introduces a synthetic dataset of 15,000 block‑stacking problems designed to improve spatial intelligence in Large Vision‑Language Models (LVLMs). The dataset covers 3D‑to‑2D projection, viewpoint transformation, and structural combination, and uses controlled color modulation to encourage anchor‑based reasoning. Experiments show that LVLMs trained on SpatialBlock outperform baselines and generalize to real‑world spatial tasks, despite the dataset’s synthetic and compact nature.
By Soohyun Ryu, Sohee Kim, Eunho Yang
arXiv:2601.22475v2 Announce Type: replace
Abstract: Building a generalist robot policy requires continuously integrating new skills while preserving previously acquired behaviors. Directly optimizing...
By Qijun He, Yuxuan Li, Mingqi Yuan, Xiaoquan Sun, Wen-Tse Chen, Jeff Schneider, Jiayu Chen
arXiv:2609.05516v1 Announce Type: cross
Abstract: Unified perception enables autonomous driving systems to perform object detection, drivable-area segmentation, and lane segmentation within a single...
By Zhiyuan Nie, Zixi Zhou, Xianbin Gu
arXiv:2511.09057v4 Announce Type: replace-cross
Abstract: A world model is a cognitive simulator of the real-world environment allowing biological agents to reason about how the world evolves, whethe...
By PAN Team, Zihan Liu, Yi Gu, Mingkai Deng, Guangyi Liu, Zeyu Feng, Qiyue Gao, Yiyan Hu, Benhao Huang, Yichi Yang, Kun Zhou, Jiannan Xiang, Zhiting Hu, Zhengzhong Liu, Eric P. Xing
arXiv:2609.06880v1 Announce Type: cross
Abstract: Reasoning over language instructions in embodied tasks such as robotics often requires understanding spatial relations from a speaker's situated pers...
By Mimo Shirasaka, Haochen Zhang, Yonatan Bisk
arXiv:2609.08153v1 Announce Type: new
Abstract: Generative diffusion models have emerged as a class of powerful techniques for various imaging applications, including but not limited to synthesis, re...
By Nian Wu, Nivetha Jayakumar, Jiarui Xing, Miaomiao Zhang
arXiv:2609.07738v1 Announce Type: cross
Abstract: LiDAR-based 3D Single Object Tracking (3D SOT) is critical for robotic perception and navigation and aims to localize dynamic objects across frames i...
By Zhaofeng Hu, Sifan Zhou, Jiahao Nie, Ziyu Zhao, Weizi Li, Ci-jyun Liang
FPicker is a topology-guided framework for filament tracing in low‑signal Cryo‑EM images. It combines a center‑endpoint representation with an open‑curve evolution module to model non‑cyclic connectivity, overcoming limitations of pixel‑wise segmenters, box‑based detectors, sequential trackers, and traditional active contours. On simulated benchmarks, FPicker improves mean spatio‑angular precision by over 40% and reduces topological gap rates by more than 60% under extreme noise, and it achieves state‑of‑the‑art performance on real EMPIAR data after fine‑tuning.
By Tingyin Zhao, Mingtao Huang, Yuan Shen
The paper introduces Reward Ensemble under Confidence (REC), a probabilistic reward learning framework for preference-based reinforcement learning that models per‑timestep reward uncertainty using an ensemble of distributional reward models. REC incorporates uncertainty into the preference loss and uses model disagreement to drive exploration, achieving 88.4% of shaped‑reward performance on acrobatic quadrotor control versus 55.2% with standard Preference PPO. The authors train policies in simulation and transfer them zero‑shot to real quadrotors, demonstrating complex acrobatic maneuvers learned solely from human preference feedback, and validate REC on a continuous‑control benchmark.
By Colin Merk, Ismail Geles, Jiaxu Xing, Angel Romero, Giorgia Ramponi, Davide Scaramuzza
arXiv:2605.10201v3 Announce Type: replace-cross
Abstract: Generalizable manipulation involving cross-type object interactions is a critical yet challenging capability in robotics. To reliably accompl...
By Zhenhao Shen, Zeming Yang, Yue Chen, Yuran Wang, Shengqiang Xu, Mingleyang Li, Hao Dong, Ruihai Wu
arXiv:2507.06625v4 Announce Type: replace-cross
Abstract: Model Predictive Control (MPC) enables reliable trajectory optimization under dynamics constraints, but often depends on accurate dynamics mo...
By Shizhe Cai, Zeya Yin, Jayadeep Jacob, Fabio Ramos
arXiv:2605.12957v2 Announce Type: replace
Abstract: Recent developments in generative models and large-scale datasets have substantially advanced 3D world generation, facilitating a broad range of do...
By Hanxin Zhu, Cong Wang, Peiyan Tu, Jiayi Luo, Tianyu He, Xin Jin, Zhibo Chen
The paper demonstrates that image knowledge distillation can be backdoored even when the teacher model is clean, by poisoning the distillation dataset with triggered and manipulated images that the teacher already classifies as a target label. The attack, effective at poisoning rates as low as 10%, uses targeted adversarial perturbations and GAN-based class transitions to embed a backdoor into the student model while preserving its performance on clean data. The study highlights that the security of knowledge distillation depends not only on the teacher but also on the integrity of the distillation data.
By Qian Ma, Chen Wu, Prasenjit Mitra, Sencun Zhu
arXiv:2609.07126v1 Announce Type: cross
Abstract: In world model planning, sensing inputs pass through an encoder and predictor before affecting planner decisions, so final task success alone cannot...
By Geonmyeong Lee, Byoung-Tak Zhang