The paper presents a benchmark to test whether vision‑language models can produce plant simulation configurations from images using in‑context learning. It focuses on cowpea plot reconstruction, requiring the models to output structured JSON that includes field and plant details. Open‑source multimodal models from the Gemma 4 and Qwen3.5 families are evaluated on synthetic and real drone datasets, using five in‑context methods, and the results show that while VLMs can generate valid JSON and estimate key agronomic metrics, their performance varies and often lags behind dataset baselines.
By Heesup Yun, Isaac Kazuo Uyehara, Earl Ranario, Lars Lundqvist, Christine H. Diepenbrock, Brian N. Bailey, J. Mason Earles
MessyKitchens introduces a new dataset of cluttered real-world kitchen scenes with detailed 3D object shapes, poses, and accurate contact information. The authors extend the SAM 3D single-object reconstruction method with a Multi-Object Decoder (MOD) to jointly reconstruct entire scenes, achieving better registration accuracy and reduced inter-object penetration compared to prior work. The dataset, benchmark, code, and pretrained models will be publicly released on the project website.
By Junaid Ahmed Ansari, Ran Ding, Fabio Pizzati, Ivan Laptev
BIDETA is a gradient‑free framework that adapts pretrained tactile models to new sensors using only a few labeled target contacts. It preserves pretrained representations while repairing sensor‑dependent feature neighborhoods through rapid support memory, support‑conditioned spectral graphs, and reliability‑gated recurrence. Experiments on multiple datasets show that BIDETA dramatically improves accuracy and speeds up adaptation compared to prior methods.
By Boheng Liu, Lan Wei, Ziyu Li, Chenghua Duan, Qing Li, Dandan Zhang, Xia Wu
arXiv:2609.28239v1 Announce Type: new
Abstract: With the development of applications like autonomous driving, object detection has gained significant attention, while also highlighting critical vulne...
By Li Zeng, Mingcheng Duan, Longfei Fan, Hangtao Zhang, Xianlong Wang, Yanchun Li, Xia Wen, Leo Yu Zhang
arXiv:2609.28312v1 Announce Type: cross
Abstract: We present VGM-VS, a visual servoing method built on a pretrained feed-forward visual geometry model. Given the current view and a reference image ca...
By Yimin Pan, Sen Wang, You Zhou, Jianfeng Gao, Pengbo Sun, Ahmed M. Naguib, Zoltan-Csaba Marton
The paper proposes a method to train efficient multi‑task manipulation policies by distilling knowledge from single‑task Conditional Flow Matching (CFM) experts. Instead of training separate models for each task, the authors transfer the experts’ learned velocity fields into a shared policy, combining this distillation signal with the original CFM objective. Experiments on RLBench demonstrate that this approach improves multi‑task performance while keeping the model size fixed, avoiding the need for larger capacity or performance drops seen with naive concatenated training.
By Shreya Deshmukh, Imen Mahdi, Nick Heppert, Abhinav Valada
arXiv:2609.27455v1 Announce Type: new
Abstract: World Action Models (WAMs) jointly model action generation and environment dynamics and are mostly built on pretrained Video Diffusion Models (VDMs). I...
By Xueji Fang, Boqiang Duan, Hua Wu, Jingdong Wang, Guo-Jun Qi
arXiv:2309.10164v3 Announce Type: replace-cross
Abstract: We develop a decentralized Perception-Action-Communication (PAC) system for multi-robot teams that enables them to collaborate in large scale...
By Saurav Agarwal, Frederic Vatnsdal, Romina Garcia Camargo, Carlos Nieto-Granda, Vijay Kumar, Alejandro Ribeiro
arXiv:2609.27378v1 Announce Type: cross
Abstract: End-to-end speech-to-speech dialogue models listen and speak simultaneously, so a continuously open acoustic channel is exposed to adversarial manipu...
By Kian Shamsaie, Iman Modarressi
arXiv:2609.27450v1 Announce Type: cross
Abstract: Vision-language-action (VLA) models handle long-horizon manipulation, yet success hinges on a few precision-critical phases where millimeter-scale er...
By Weihui Zhao, Xiaohan Yan, Zunian Wan, Xuan Du, Zhaozhan Chi, Jianbo Mao, Ruipu Wu, Rushuai Yang, Houlin Li, Shukai Yang, Jing Wu, Yuxiang Yan, Yongcheng Liu, Chuankang Li, Guanghui Ren, Wei Shan, Maoqing Yao
arXiv:2609.27656v1 Announce Type: cross
Abstract: Physical intelligence requires more than predicting how the world may evolve: predictions must remain actionable as the world continues to change. We...
By Jisong Cai, Yao Mu, Ganlin Yang, Zhe Cao, Zhangzheng Tu, Xing Gao, Kailin Li, Xinyu Zhan, Lixin Yang, Yangkun Zhu, Haoxiang Ma, Ming Zhou, Qiaojun Yu, Yufei Xue, Liqun He, Yifei Yao, Yifan Zhu, Long Ling, Bingqi Jiang, Haoyu Guo, Xueyue Zhu, Bowen Zhou, Bin Zhao, Tianfan Xue, Chunhua Shen, Weinan Zhang
arXiv:2609.27816v1 Announce Type: cross
Abstract: Safe coordination in heterogeneous machine-to-machine (M2M) robotic systems is challenging when robots differ in sensing capabilities, environmental...
By Mohamed Dwedar, Ahmad Hafez, Alexander Jesser, Amr Alanwar
arXiv:2603.15220v2 Announce Type: replace
Abstract: Strict anonymity of model responses is a key for the reliability of voting-based leaderboards, such as LM Arena. While prior studies have attempted...
By Minsung Cho, Jaehyung Kim
arXiv:2609.27076v1 Announce Type: new
Abstract: Open-vocabulary visual grounding enables robots to localise task-relevant entities from natural-language queries without dependence on predefined perce...
By Linus Nwankwo, Muslim Alaran, Christian Rauch, Stanley Chukwuebuka Obilikpa, Elmar Rueckert
arXiv:2609.27227v1 Announce Type: new
Abstract: Objective assessment of robotic surgery uses instrument kinematics, which must be reconstructed when only video is available. We introduce a kinematic...
By Mehmet Kerem Turkcan, Soham Samal, Zoran Kostic
arXiv:2609.28236v1 Announce Type: new
Abstract: Long-horizon embodied interaction requires agents to retain and continually update information about the environment as they observe, act, and encounte...
By Lizhou Liang, Xinyu Zhong, Miao Pan, Xiaohe Zhou, Xuanyu Liu, Qinfeng Li, Peng Li, Jintao Chen, Xuhong Zhang, Wenqi Zhang
arXiv:2609.27094v1 Announce Type: cross
Abstract: Automatic tagging is a core task in Music Information Retrieval (MIR), yet most tagging systems exploit only audio. Live music performance is inheren...
By Alexandros Alexiou, Charilaos Papaioannou, Alexandros Potamianos
arXiv:2604.15221v3 Announce Type: replace-cross
Abstract: Safe human-robot collaboration (HRC) requires accurate human pose estimation and motion prediction to prevent critical collisions. Existing c...
By Jakob Thumm, Marian Frei, Tianle Ni, Matthias Althoff, Marco Pavone
arXiv:2609.23407v2 Announce Type: replace-cross
Abstract: Humans can effortlessly localize the direction of a sound source and integrate it with visual cues for reasoning, yet this remains challengin...
By Ruixun Liu, Yuxuan Wang, Jiacheng Xie, Yuhuan You, Donghua Cai, Junming Lin, Xiong-Hui Chen, Zhifang Guo, Yunfei Chu, Qize Yang, Xize Cheng, Jin Xu, Yiwu Zhong