The paper introduces a cross‑modal pseudo‑labeling pipeline for unsupervised domain adaptation in semantic segmentation, particularly for waste sorting. It combines SAM for class‑agnostic region proposals with EVA‑CLIP to assign semantic labels via region‑text similarity, applying confidence filtering to ensure reliable pseudo‑labels for self‑training. An optional BLIP‑based language‑grounded verification further refines ambiguous regions, and the method shows consistent improvements over source‑only baselines on synthetic‑to‑real driving and lab‑to‑factory waste sorting shifts.
By Udo Schlegel, Shubhangi, Gabriel Dax, Sai Rahul Kaminwar, Florian Karl, Thomas Seidl
arXiv:2609.01351v1 Announce Type: cross
Abstract: Online planning under uncertainty remains a fundamental challenge for robotic systems operating in partially observable environments with high-dimens...
By Jiho Lee, Nisar Ahmed, Kyle Hollins Wray, Zachary Sunberg
The paper introduces DroneCATS-Agent, a modular framework that places a multimodal large language model (MLLM) at the core of a drone’s control loop, allowing the model to decide actions solely from prompts. It presents the DroneCATS benchmark, evaluating MLLMs on four tasks—approaching, tracking, searching, and multi‑drone commanding—without fine‑tuning or function‑calling. Results show that while small open models can navigate reliably, they often fail by mismanaging protocol termination, highlighting a gap between perception and action planning in current MLLMs.
By Jaewoo Park, Minyoung Lee, Sukmin Seo, Moonbin Yim, Hyunwook Yoon, Dohoon Ryu, Daehee Kim, Myungseo Song, Jihyuk Byun, Seunggyu Chang, Taeho Kil, Jiseob Kim, Bado Lee, Geewook Kim
The paper examines how responsibility is assigned when AI systems fail, proposing a sociotechnical theory that distinguishes between AI incidents, organisational crises, and scandals. It argues that the configuration of an incident shapes actor-specific attribution, which in turn influences perceptions of capability, integrity, fairness, and relationships, and that public moralisation can elevate an incident to scandal. The authors introduce ‘accountable transparency’—a response framework combining timely notice, intelligible accounts, role acknowledgement, remedy, evidence of correction, and recourse—as a way to manage blame, trust, and communication credibility.
By Mohammad Saleh Torkestani, Taha Mansouri
arXiv:2609.01556v1 Announce Type: cross
Abstract: We evaluate embedding retrieval where surface form and meaning are pulled apart on purpose: retrieving items that share underlying structure but not...
By Nabira Rashid, Manolis Kellis
EEG-VID is a task‑guided latent predictive pretraining framework designed to improve EEG decoding across session and subject shifts. It predicts future latent EEG states from recent history using an exponential‑moving‑average target encoder and weak task guidance, then fine‑tunes with supervised learning. The method achieves significant accuracy gains on VIG‑48 and BCI Competition datasets, and demonstrates effective assistive target selection in a robot‑scene study.
By Guanzhong Sun, Junyi Ma, Yuxuan Wu, Yanzi Miao
arXiv:2609.00474v1 Announce Type: cross
Abstract: LLMs are increasingly deployed as orchestrators that coordinate specialized subagents to solve complex tasks through natural language. However, in ma...
By Harini S I, Somesh Singh, Yaman K Singla, Rajiv Ratn Shah, David Doermann, Balaji Krishnamurthy
SEBA is a sample‑efficient framework for black‑box adversarial attacks on visual reinforcement learning agents. It combines a shadow Q model, a generative adversarial network for imperceptible perturbations, and a world model to simulate dynamics, reducing real‑world queries. Experiments on MuJoCo and Atari show SEBA significantly lowers cumulative rewards while preserving visual fidelity and requiring far fewer environment interactions than previous methods.
By Tairan Huang, Yulin Jin, Junxu Liu, Qingqing Ye, Haibo Hu
arXiv:2609.01120v1 Announce Type: cross
Abstract: Early recognition of lane-change intention is essential for proactive decision-making in autonomous driving and advanced driver assistance systems. T...
By Woong-Chan Byun, Seung-Hyun Kong
Hydra introduces a marker‑free RGB‑D hand‑eye calibration method that leverages a novel ICP algorithm with a robust point‑to‑plane objective on a Lie algebra. Experiments on three serial manipulators and two RGB‑D cameras show that with only three random robot configurations the method achieves about 90% successful calibrations, 2–3× faster convergence to the global optimum, and 2 orders of magnitude faster convergence time (0.8 ± 0.4 s) compared to other marker‑free baselines. The approach delivers improved accuracy (5 mm in task space versus 7 mm for classical methods) while remaining marker‑free, and the authors provide an open‑source dataset, code, and ROS 2 integration.
By Martin Huber, Huanyu Tian, Christopher E. Mower, Lucas-Raphael M\"uller, S\'ebastien Ourselin, Christos Bergeles, Tom Vercauteren
arXiv:2505.13180v3 Announce Type: replace
Abstract: Integrating Large Language Models with symbolic planners is a promising direction for obtaining verifiable and grounded plans, with recent works ex...
By Matteo Merler, Nicola Dainese, Minttu Alakuijala, Giovanni Bonetta, Pietro Ferrazzi, Yu Tian, Bernardo Magnini, Pekka Marttinen
ZimaBlue is a scalable framework that learns generalizable World Action Models (WAMs) from large-scale egocentric videos. It follows a three-stage curriculum: causal video pre‑training, video‑action mid‑training with a unified action representation, and final specialization to a target robot. The system employs an asynchronous Slow‑Fast architecture to enable real‑time 30 Hz action prediction, achieving a jump in real‑robot zero‑shot success from 36.1% to 77.8% when leveraging over 120,000 hours of embodied video.
By Xionghao Wu, Yijun Yang, Shiyang Zhou, Haoze Sun, Jianhui Liu, Songsong Yu, Jiyao Zhang, Wenbo Li, Bo Wang, Guoqing Ma, Lin Song, Renjie Liao, Shenghe Zheng, Wei Tang, Xiaojuan Qi, Yanwei Li, Yuan Zhang, Zhuotao Tian, Haoyang Huang, Nan Duan
arXiv:2609.01418v1 Announce Type: cross
Abstract: To mitigate the sample complexity of real-world reinforcement learning (RL), a common practice is to first train a policy in a simulator, where sampl...
By Tingting Ni, Maryam Kamgarpour
The paper introduces Hand2Bot, an RGB‑D video dataset designed for human‑to‑robot handover scenarios, capturing body posture and facial expressions amid real‑world noise. It also proposes PassGen, a generative pipeline using stable video diffusion and an Intention‑Aware Temporal Face Encoder to synthesize realistic handover sequences while maintaining hand‑object consistency. A morphology‑based depth editing strategy is employed to replicate realistic sensor noise, and experiments show that training on PassGen yields high intention identification accuracy, low false trigger rates, and robust zero‑shot transfer to a physical robot platform.
By Tianyu Sun, Zhoujie Fu, Zihui Gao, Bang Zhang, Guosheng Lin
arXiv:2507.20881v3 Announce Type: replace
Abstract: Endoscopic depth estimation is a critical technology for improving the safety and precision of minimally invasive surgery. It has attracted conside...
By Ke Niu, Zeyun Liu, Xue Feng, Heng Li, Naian Xiao, Binghua Su, Qika Lin, Kaize Shi
Real-world robotic assembly at sub-millimeter tolerances demands spatial precision, compliant interaction, and robustness to contact failures. We present Facet-0, a robotic foundation model that predi...
arXiv:2608.30935v1 Announce Type: cross
Abstract: Embodied navigation requires agents to translate heterogeneous goals and visual observations into actions across tasks, environments, and robot embod...
By Shaoan Wang, Aocheng Luo, Fei Huang, Jingyi Xu, Xiaoyang Wang, Yueyu Wang, Qianli Ma, Fan Yang, Ran Mei, Jia Wei, Jiangpeng Hu, Xuhao Liu, Hongming Chen, Yuanbin Shao, Yiyang Lin, Ziliang Li, Liang Pan, Xinhang Liu, Yuntao Ma, Tingxiang Fan
The paper introduces NavMCP, a scaffolding framework that couples vision‑language models (VLMs) with navigation foundation models (NFMs) to enable long‑horizon physical‑world agents. NavMCP orchestrates three communication channels—intent, observation, and memory—to allow the VLM to decide what evidence to seek and the NFM to ground semantic sub‑goals into closed‑loop navigation, without retraining either model. The approach achieves state‑of‑the‑art results on several embodied question‑answering benchmarks and significantly outperforms episodic interfaces on the Unitree Go2 robot as task horizons lengthen.
By Zixing Lei, Gengze Zhou, Xiong-Hui Chen, Jiazhao Zhang, Yiyang Huang, Hang Yin, Haoqi Yuan, Qi Wu, Weixin Li, Siheng Chen
arXiv:2604.22230v2 Announce Type: replace-cross
Abstract: Performance manipulation arises when agents exploit easily measurable, routine tasks to inflate observable outcomes without contributing genu...
By Xiaoyun Qiu, Yang Yu, Haifeng Xu
The paper introduces the concept of intervention fidelity in latent world models, measuring whether a model’s open‑loop transitions align with actual environment interventions. Experiments on TD‑MPC2, Cheetah, and DreamerV3 show that high reward fit does not guarantee fidelity, and that self‑supervised models can outperform task‑anchored ones in preserving intervention effects. The authors propose a capture‑gated audit to localize failures and argue that fidelity must be directly audited on the model’s native interface.
By Donna Vakalis