arXiv:2609.08402v1 Announce Type: cross
Abstract: Air-Ground Object Search (AGOS) in urban environments is a challenging embodied task, which requires an Unmanned Aerial Vehicle (UAV) and an Unmanned...
By Boao Yu, Zimo Chen, Junreng Rao, Yue Hu, Zhengqiu Zhu, Yong Zhao, Rusheng Ju
The paper demonstrates that image knowledge distillation can be backdoored even when the teacher model is clean, by poisoning the distillation dataset with triggered and manipulated images that the teacher already classifies as a target label. The attack, effective at poisoning rates as low as 10%, uses targeted adversarial perturbations and GAN-based class transitions to embed a backdoor into the student model while preserving its performance on clean data. The study highlights that the security of knowledge distillation depends not only on the teacher but also on the integrity of the distillation data.
By Qian Ma, Chen Wu, Prasenjit Mitra, Sencun Zhu
arXiv:2609.07511v1 Announce Type: cross
Abstract: Autonomous driving relies on High Definition (HD) maps for safe navigation. Traditional HD maps construction is costly in hardware, data and human re...
By Clara Gomez, Alberto Jaenal, Antonio Artu\~nedo, Jorge Godoy, Jorge Villagra
arXiv:2605.12957v2 Announce Type: replace
Abstract: Recent developments in generative models and large-scale datasets have substantially advanced 3D world generation, facilitating a broad range of do...
By Hanxin Zhu, Cong Wang, Peiyan Tu, Jiayi Luo, Tianyu He, Xin Jin, Zhibo Chen
arXiv:2609.05593v1 Announce Type: cross
Abstract: Autonomous robot navigation failures differ not only in categorical severity but also in the physical context in which they occur. A near-miss at low...
By Rifa Ferzana
arXiv:2609.07534v1 Announce Type: cross
Abstract: Contact-rich assembly remains challenging because it requires submillimeter spatial accuracy and reliable interpretation of forces during sustained c...
By Yuhan Wang, Yurou Chen, Hongye Jiang, Wenzhao Lian
CogniDir is an adaptive distributional learning framework designed to improve fake news detection against new psychologically grounded malicious comments generated by Large Language Models. It reframes robust detection as a dynamic data mixture optimization problem, using cognitive psychology to formalize adversarial paradigms and an information‑theoretic score to guide adaptive sampling of training data. Experiments on three benchmarks show that CogniDir achieves state‑of‑the‑art robustness, boosting F1 scores by up to 17.9% over existing baselines under heterogeneous AI‑generated attacks.
By Zhao Tong, Chunlin Gong, Yimeng Gu, Haichao Shi, Qiang Liu, Shu Wu, Xingcheng Xu, Xiao-Yu Zhang
arXiv:2609.08273v1 Announce Type: new
Abstract: Agent memory systems have demonstrated significant potential in long-term dialogue, personalized assistants, and video understanding. However, continuo...
By Junxi Wang, Te Sun, Jiayi Zhu, Chen Zhang, Siyuan Li, Xuyang Liu, Zichen Wen, Xiaobing Tu, Jinkui Ren, Xiantao Zhang, Ziqi Yuan, Linfeng Zhang
arXiv:2609.07047v1 Announce Type: cross
Abstract: Robotic manipulation often requires acting on information that is no longer visible, yet Vision-Language-Action policies are usually evaluated when t...
By Haiyang Sun, Haoxiao Wang, Junming Chen, Weicheng Fang, Zihao Su, Jingkun Yi, Wenyou Yi, Hao Chen, Zhou Zhao
CALIPER is a new benchmark that tests whether pretrained visual encoders can infer physical properties such as mass and friction from images. The test involves striking an object twice at known speeds, showing a third strike only up to contact, and asking a linear readout on frozen features to predict how far the object slides. Results show that in clean, fixed‑camera scenes all representations perform similarly, but when camera, lighting, and clutter are varied, only encoders that truly infer physics—like V‑JEPA 2—maintain performance, while random or raw pixel representations fail.
By Aman Mehta, Riya Baviskar
The paper introduces 3DWay, a method that predicts 3D-consistent waypoints for robot manipulation by first generating multi‑view consistent 2D waypoints and then triangulating them. This approach addresses the 3D ambiguity inherent in 2D trajectory predictions and leverages pretrained vision‑language models to provide explicit 3D motion specifications. Experiments demonstrate that 3DWay improves 3D spatial grounding and vision‑language reasoning, enhancing generalization for robot manipulation tasks.
By Ziqin Huang, Yingyue Li, Chenyangguang Zhang, Ruida Zhang, Yuxin Chen, Gu Wang, Xingyu Liu, Masayoshi Tomizuka, Xiangyang Ji
ControlTac is a two‑stage framework that generates realistic tactile images conditioned on a single reference image, contact force, and contact pose. By incorporating these physical priors, it produces realistic samples across different sensors and captures task‑relevant variations. Experiments in object insertion, imitation learning, and object weighting show that datasets augmented with ControlTac consistently improve performance in dynamic real‑world settings.
By Dongyu Luo, Kelin Yu, Amir-Hossein Shahidzadeh, Cornelia Ferm\"uller, Yiannis Aloimonos, Ruohan Gao
arXiv:2609.07713v1 Announce Type: new
Abstract: Generative and agentic AI are reshaping both the production and evaluation of scientific research. These developments are often studied separately, as...
By Chenguang Wang, Ming Li, Adebayo Braimah, Chenrui Fan, Tuo Wang, Weijie Guan, Ruiyi Zhang, Tianyi Zhou, Dawei Zhou
arXiv:2609.07559v1 Announce Type: new
Abstract: How do you validate a cheap, deterministic proxy for an oracle that is expensive, rate-limited, and non-stationary? We present a protocol built on adve...
By Elisha Bajemon, Andre-Louis Rochet
arXiv:2609.07111v1 Announce Type: cross
Abstract: Quadruped robot locomotion policies are often trained using reinforcement learning, which in turn relies heavily on hand-crafted reward functions. De...
By Merve Atasever, Keyan Azbijari, Cagan Bakirci, Alfredo Reina Corona, Tolga Izdas, Richard Yang, Erdem Biyik, Jyotirmoy V. Deshmukh
The paper introduces Distributed Dexterous Manipulation (DDM), a challenging control problem involving 64 soft delta robots arranged in an 8x8 grid. It presents a framework using spatially conditioned Multi-Agent Transformers (MATs) with adaptive layer norm, spatial contrastive embeddings, and a behavior cloning method fine‑tuned by Soft Actor Critic. Experiments demonstrate that MATs refine actions through stacked attention blocks, enabling long‑horizon planar manipulation in simulation and real‑world settings, while an action‑selection strategy reduces robot usage by about 65% and lowers wear‑and‑tear, achieving an average error of ~1.5 cm.
By Sarvesh Patil
arXiv:2609.07998v1 Announce Type: new
Abstract: We study the control of Markov decision processes in which the quality of a policy is evaluated by a dynamic, time-consistent Markov risk measure rathe...
By Aayush Patel, Andrzej Ruszczy\'nski
arXiv:2609.10464v1 Announce Type: cross
Abstract: Joint-Embedding Predictive Architecture (JEPA) world models learn a compact latent representation of the world that supports prediction and planning,...
By Andy Zeyi Liu, Haoran Sun, Lucas Baker, Randall Balestriero, John Sous
arXiv:2609.10522v1 Announce Type: cross
Abstract: Foundation vision-language models (VLMs) exhibit broad intelligence about the world, yet translating this intelligence into robot control remains cha...
By Yanzhe Chen, Zechen Bai, Zhijun Cao, Wenzheng Zeng, Kevin Qinghong Lin, Yiqi Lin, Guoqiang Liang, Kevin Yuchen Ma, Qiming Huang, Mike Zheng Shou
arXiv:2609.07738v1 Announce Type: cross
Abstract: LiDAR-based 3D Single Object Tracking (3D SOT) is critical for robotic perception and navigation and aims to localize dynamic objects across frames i...
By Zhaofeng Hu, Sifan Zhou, Jiahao Nie, Ziyu Zhao, Weizi Li, Ci-jyun Liang