arXiv:2608.29967v1 Announce Type: cross
Abstract: Vision-Language-Action (VLA) models demonstrate strong semantic understanding yet exhibit systematic failures during deployment. The conditions under...
By Owen Kwon, Pablo Ortega-Kral, Arthur Bucker, Jean Oh
arXiv:2608.29537v1 Announce Type: cross
Abstract: Frozen vision-language-action (VLA) policies offer broad manipulation skills but execute open-loop action chunks without tracking task progress, so t...
By Hongbo Gao, Zeyu Ni, Xin Wen, Siyu Xu, Ruifeng Li
arXiv:2608.29114v1 Announce Type: cross
Abstract: Vision-and-Language Navigation (VLN) requires agents to reason over accumulated observations while continuously exploring unseen regions. However, ex...
By Yuxiang Xiao, Xibei Chen, Xin Zhou, Jie Chen, Yifeng Zhang, Guillaume Sartoretti
arXiv:2608.29937v1 Announce Type: new
Abstract: Latent world-action models avoid rendering future pixels by predicting an action-relevant visual subgoal in feature space. LaWAM established this formu...
By Yafei Zhang, Nan Wu
arXiv:2608.29315v1 Announce Type: cross
Abstract: This work introduces Semantically-Guided Exploration (SGE), a modular exploration framework for ground vehicles that integrates pixel-level semantic...
By Christopher Tatsch, Yu Gu
arXiv:2608.30821v1 Announce Type: cross
Abstract: Composable scene modeling aims to recover a real indoor scene as complete, editable object assets arranged as observed, giving robot simulation and e...
By Minghan Qin, Yuang Wang, Xiuyu Yang, Yushi Long, Yujian Zhang, Ruihuan Wang, Kai Ye, Yangang Zhang, Hang Li
arXiv:2608.28733v1 Announce Type: cross
Abstract: Indoor 3D Scene Graphs (3DSGs) represent environments as multi-layer hierarchies that connect observed geometric primitives (e.g., planes) to higher-...
By Jose Andres Millan-Romera, Samuel Cognolato, Holger Voos, Jose Luis Sanchez-Lopez, Luciano Serafini
arXiv:2608.28948v1 Announce Type: cross
Abstract: Over the course of the last decade, neural networks have grown from an academic curiosity to moving the markets of nations. Despite this explosion in...
By David Aram Yunis
arXiv:2608.30897v1 Announce Type: new
Abstract: World models are becoming core infrastructure for embodied intelligence, with action-conditioned video generation providing controllable predictions of...
By Jianjie Fang, Xvyuan Liu, Ziyou Wang, Rongze Tang, Zhaolu Wang, Zhuohang Li, Xin Zhang, Haisheng Su, Chen Gao, Wei Wu, Xinlei Chen, Yong Li
arXiv:2608.29426v1 Announce Type: cross
Abstract: Reliable semantic representations derived from city-scale 3D models are increasingly important for urban analysis, infrastructure monitoring, autonom...
By Alexander Rusnak, Sophia Kovalenko, Jingru Wang, Ismail Moudden, Xiru Wang, Fr\'ed\'eric Kaplan
arXiv:2608.29475v1 Announce Type: new
Abstract: Surface material recognition from incomplete visual observations remains a challenging problem in robotic perception and environmental understanding. T...
By Sindhuja Penchala, Sudip Mittal, Noorbakhsh Amiri Golilarz
arXiv:2608.29601v1 Announce Type: cross
Abstract: We present $\mathcal{N}_0$-Foundation, a paradigm for tactile-enabled embodied manipulation, which integrates tactile sensing hardware, large-scale m...
By NeoteAI Team, Fudan TEAI Team
arXiv:2608.29208v1 Announce Type: cross
Abstract: Vision-Language-Action (VLA) models, built upon Vision-Language Models (VLMs), have significantly enhanced robotic capabilities by leveraging interne...
By Sunghwan Han, Youngtae Han, Youngmin Yi
arXiv:2608.29459v1 Announce Type: new
Abstract: Skills, as a useful abstraction for the procedural capabilities of large language models (LLMs), capture how models perform structured, multi-step reas...
By Xunyi Jiang, Junda Wu, Yuxin Xiong, Sheldon Yu, Tong Yu, David Arbour, Ritwik Sinha, Julian McAuley, Hongyi Wen
arXiv:2608.29596v1 Announce Type: new
Abstract: Autonomous large language model (LLM) agents increasingly face reliability, context consumption, and execution stability bottlenecks when deployed on c...
By Sanket Badhe, Deep Shah, Priyanka Tiwari, Nehal Kathrotia
arXiv:2607.15542v2 Announce Type: replace
Abstract: On-the-fly reconstruction is a key requirement for many applications in robotics and autonomous navigation. Variational Bayes Gaussian Splatting (V...
By Damani Mguni-Coker
arXiv:2608.31029v1 Announce Type: cross
Abstract: End-to-end autonomous driving models plan future trajectories from raw sensor input. While earlier driving benchmarks often measured deviation from t...
By Christian L\"owens, Thorben Funke, Alexandru Paul Condurache
arXiv:2608.30804v1 Announce Type: new
Abstract: Monitoring the health of heterogeneous industrial robot fleets is severely challenged by the multi-modal nature of their operational cycles and a persi...
By Martin Bonsergent-Brachet, Jesse Read, Dany Abboud
arXiv:2608.28778v1 Announce Type: cross
Abstract: Autonomous vehicles (AVs) rely on accurate camera-LiDAR calibration for multimodal sensor fusion. In practice, calibration can drift due to vibration...
By Liangkai Liu, Qingzhao Zhang, Kang G. Shin
arXiv:2608.29923v1 Announce Type: cross
Abstract: Open-vocabulary semantic segmentation (OVSS) relies on vision-language alignment to recognize arbitrary text-defined categories, yet this alignment i...
By Chandler Timm C. Doloriel, Yunbei Zhang, Sarthak Kumar Maharana, Muhammad Salman Siddiqui, Tor Kristian Stevik, Fadi Al Machot, Kristian Hovde Liland, Habib Ullah