arXiv:2511.09057v4 Announce Type: replace-cross
Abstract: A world model is a cognitive simulator of the real-world environment allowing biological agents to reason about how the world evolves, whethe...
By PAN Team, Zihan Liu, Yi Gu, Mingkai Deng, Guangyi Liu, Zeyu Feng, Qiyue Gao, Yiyan Hu, Benhao Huang, Yichi Yang, Kun Zhou, Jiannan Xiang, Zhiting Hu, Zhengzhong Liu, Eric P. Xing
arXiv:2609.05516v1 Announce Type: cross
Abstract: Unified perception enables autonomous driving systems to perform object detection, drivable-area segmentation, and lane segmentation within a single...
By Zhiyuan Nie, Zixi Zhou, Xianbin Gu
arXiv:2609.07047v1 Announce Type: cross
Abstract: Robotic manipulation often requires acting on information that is no longer visible, yet Vision-Language-Action policies are usually evaluated when t...
By Haiyang Sun, Haoxiao Wang, Junming Chen, Weicheng Fang, Zihao Su, Jingkun Yi, Wenyou Yi, Hao Chen, Zhou Zhao
arXiv:2609.06079v1 Announce Type: new
Abstract: Vision-Language-Action (VLA) policies leverage pretrained vision-language models (VLMs) to guide action generation for robot control. VLMs provide hier...
By Zheng Lu, Haoran Liao, Wanqi Zhong, Yunhe Ni, Lijie Wang, Xingjie Fan, Zhisheng Chen, Yantang Qu, Meijia Chen, Tianyu Xin, Zirui Song, Yiming Li
arXiv:2609.10464v1 Announce Type: cross
Abstract: Joint-Embedding Predictive Architecture (JEPA) world models learn a compact latent representation of the world that supports prediction and planning,...
By Andy Zeyi Liu, Haoran Sun, Lucas Baker, Randall Balestriero, John Sous
arXiv:2609.08602v1 Announce Type: new
Abstract: Embodied planning increasingly relies on vision-language models (VLMs) to translate instructions and visual observations into executable action sequenc...
By Tianyi Ma, Parisa Kordjamshidi
arXiv:2609.07741v1 Announce Type: new
Abstract: Persistent AI assistants are intended to extend human attention, memory, and coordination across changing digital and physical environments. To be trul...
By Jo\~ao Dias Ferreira
arXiv:2609.10506v1 Announce Type: cross
Abstract: Action-conditioned latent world models predict future visual representations, enabling zero-shot goal-conditioned robot planning and control. However...
By Nisarga Nilavadi, Ralf R\"omer, Moritz Reuss, Michael Krawez, Tobias J\"ulg, Angela P. Schoellig, Rudolf Lioutikov, Wolfram Burgard
arXiv:2609.05834v1 Announce Type: new
Abstract: World models promise a general route to embodied intelligence: learn predictive dynamics once, then reason, plan, and act with them. Increasingly, the...
By Todd Y. Zhou, Daniel Zhang
arXiv:2609.10082v1 Announce Type: cross
Abstract: Accurate camera intrinsic calibration is fundamental to robot perception, and the accuracy depends on the quality of the collected images. However, e...
By Xiangcheng Hu
arXiv:2609.07126v1 Announce Type: cross
Abstract: In world model planning, sensing inputs pass through an encoder and predictor before affecting planner decisions, so final task success alone cannot...
By Geonmyeong Lee, Byoung-Tak Zhang
The paper introduces BIFTA, a Brain‑Inspired Few‑Shot Tactile Adaptation framework that enables a frozen encoder to adapt quickly to an unknown tactile sensor using only a small labeled support set. It preserves pretrained representations via dual‑view statistical memory, builds support‑conditioned spectral graphs to correct sensor‑dependent feature neighborhoods, and employs uncertainty‑gated recurrent propagation to reinforce reliable cross‑query evidence. Benchmarks on three tactile datasets demonstrate that BIFTA dramatically improves adaptation performance, achieving an 87.09% mean Sparsh accuracy on SITR with just 10% labeled data—an increase of 47.22 percentage points over the best prior method.
By Boheng Liu, Ziyu Li, Xia Wu
The paper investigates adding Greek language support to a Cosmos3 vision‑language‑action robot policy using only machine‑rephrased instructions and no architectural changes. It finds that many evaluation metrics can give misleading results, and that multilingual training with Greek yields a modest 6.7‑7.1 point advantage over a control, reaching about 40% of English performance. The study also shows that overfitting to a single translator’s phrasing can be mitigated by training on multiple phrasings, while warm‑starting from a language‑adapted world model or unfreezing the text tower actually harms performance.
By Ayoub Kirouane, Georgios Giaples, Christos Petrocheilos
The paper investigates how reinforcement‑learning policies for legged robots encode gait information by examining the effective rank of the policy Jacobian conditioned on gait phase. It finds that common architectural features such as layer normalization and residual connections allocate more representational capacity to swing than stance, a pattern absent in vanilla MLPs. Leveraging these insights, the authors propose a simple recipe that improves sim‑to‑real transfer, reducing joint jitter on a physical Spot robot by roughly three‑fold.
By Felipe Tommaselli, Thiago H. Segreto, Juliano D. Negri, Ricardo V. Godoy, Marcelo Becker
The paper presents a multi‑modal deep learning model that uses temporal attention to detect internal welding defects such as porosity, lack of penetration, fusion, undercut, and cold lap in fillet joints during real‑time Gas Metal Arc Welding. Trained on images and sound data from an industrial collaborative welding robot, the model achieves an F1 score of 0.99. Explainable AI techniques are applied to interpret the model’s behavior, highlighting key image and sound spectrogram regions and the most effective modality for each defect type, thereby enhancing trust and reliability in AI‑driven welding inspection.
By Mobina Mobaraki, Mahyar Asadi, Klaske Van Heusden, Guy A. Dumont
ARC‑Bench is a new benchmark that tests whether frozen JEPA‑style latent world models can correctly rank candidate actions by latent distance. The study finds that the assumption of latent rankability fails dramatically in both navigation and manipulation tasks, with the top‑scored actions often being suboptimal. Closed‑loop replanning masks this defect, but reducing replanning frequency reveals the underlying ranking failures.
By Zhengshu Zhang, Zhiyuan Li
The paper studies how Bird's‑Eye‑View (BEV) maps predicted by Cross‑View Transformers (CVT) can be used directly as inputs to a Behavior‑Cloning (BC) driving policy in the CARLA simulator. It introduces a six‑channel BEV representation and a Kernel Density Estimation (KDE) weighting scheme to focus learning on underrepresented maneuvers. Closed‑loop tests show that the KDE‑weighted model is the only predicted‑BEV agent to finish an episode without infractions, highlighting that global segmentation scores are poor proxies for driving performance and that prediction quality at critical geometries, especially the route channel, is key to reliable navigation.
By Felipe Carlos dos Santos, Eric Antonelo, Gustavo Claudio Karl Couto
The paper presents a solution for the UCF UrbanTwin LUMPI Track in the Sim-to-Real LiDAR Challenge, where a detector trained solely on synthetic data must perform on real LiDAR frames. The approach tackles the Sim2Real gap through data alignment, diversified sampling, augmentation, and specialized detectors, followed by class-aware fusion and calibration techniques. The final submission achieved a Combined Score of 0.4692, a Detection Score of 0.1797, a Realism Score of 0.9035, and a 3D mAP@0.5 of 0.1258.
By Pu Luo, Cong Xu, Yumei Li, Kexin Zhang, Licheng Jiao, Wenping Ma, Lingling Li
The paper introduces a structure‑aware federated learning framework for segmenting catheters and guidewires in X‑ray fluoroscopy. It presents a new benchmark dataset, CathAction, and a shape‑sensitive loss that improves segmentation accuracy. The approach extends to federated learning, adding projected gradient descent for adversarial optimization, and includes a diffusion‑based synthetic data generator that boosts performance under data scarcity.
By Chayun Kongtongvattana
D3ARC is an asynchronous distributed hierarchical framework designed for time‑critical wildfire detection using multiple robotic agents. It enables cooperative perception, shared situational awareness, and coordinated actions while a remote controller asynchronously directs each robot’s motion. The system incorporates safe navigation, coverage efficiency, and a forward‑looking capability to evaluate candidate strategies before execution, achieving up to 94% mission success and 89.4% detection confidence in realistic simulations.
By Nikolaos Koursioumpas, Lina Magoula, Nancy Alonistioti, Ramin Khalili