arXiv:2610.01746v1 Announce Type: cross
Abstract: Autonomous driving systems have become a central focus of intelligent transportation research, with End-to-End Learning and Modular Architectures off...
By Kartik B. Kapse
arXiv:2610. 01834v1 Announce Type: new Abstract: Judgment models such as \jev{} return, in a single call and without reasoning text, a probability for each described option.
By Yaodong Yang, Hongyao Tang, Yi Ma, Xingyu Fan, Weixun Wang, Jinpeng Li, Tianpei Yang
arXiv:2610.02000v1 Announce Type: new
Abstract: Adverse conditions such as rain, snow, fog, and dust remain challenging for camera-based perception in autonomous driving. We study multi-class weather...
By Hossein Maghsoumi, George Atia, Yaser P. Fallah
arXiv:2609.39995v2 Announce Type: new
Abstract: Diffusion planners exhibit strong capabilities in generating multimodal trajectories. However, existing methods primarily rely on expert demonstrations...
By Jiaxi Ye, Chunji Lv, Guoren Wang, Changsheng Li
arXiv:2609.36416v2 Announce Type: replace-cross
Abstract: Robots operating in real-world environments must often execute complex, multi-step bimanual tasks over long horizons rather than single, isol...
By Jade Choghari, Pepijn Kooijmans, Mansi Agarwal, Yusuf Umut Ciftci, Aseem Doriwala, Catherine Weaver, Mouli Sivapurapu, Kai Yang, Thomas Wolf, Jackson Lee, Pragna Mannam
arXiv:2609.40245v2 Announce Type: cross
Abstract: Robot navigation in dynamic, human-centered environments requires socially-compliant decisions grounded in robust scene understanding. Recent Vision-...
By Nathan Tsoi, Michael J. Munje, Tejas Oberoi, Rishab Maheshwari, Pengen Zheng, Tanush Chauhan, Peter Stone, Joydeep Biswas
Screw Attention introduces a transformer layer that treats the relation between two bodies as a spatial transform rather than a graph edge, enabling each token to represent a body with a pose and its relative pose or joint screw. This design ensures equivariance to independent frame changes and allows a single layer to capture rigid‑body velocity recursion. Experiments on simulated manipulation tasks show that Screw Attention matches or outperforms other network architectures, achieving high success rates on LIBERO‑Spatial with far fewer parameters and maintaining performance under frame convention changes and pose noise.
By Aly Magassouba
Existing image forgery detectors often suffer from generalization to unseen manipulation methods due to the limited ability to capture transferable forensic cues. Recent cross-reconstruction based met...
The paper introduces Diffusion-Augmented Markov Decision Processes (DA‑MDPs), a framework that extends Maximum Entropy Reinforcement Learning to diffusion-based policies. DA‑MDPs treat each reverse‑diffusion step as an RL decision, deriving a tractable reverse‑KL bound that decomposes across denoising transitions and yields diffusion‑augmented soft rewards, value functions, and policy objectives. The authors implement this framework with PPO, REPPO, and a maximum‑entropy WPO variant, showing improved continuous‑control performance, higher success rates on manipulation tasks, and memory‑efficient training with action chunking.
By Sebastian Sanokowski, Kaustubh Patil, Majid Khadiv
arXiv:2605.18727v2 Announce Type: replace-cross
Abstract: Evaluating embodied systems with real dexterous hardware requires more than isolated motor-skill tests: an agent must perceive a changing sce...
By Feng Chen, Tianzhe Chu, Li Sun, Pei Zhou, Zhuxiu Xu, Shenghua Gao, Yuexiang Zhai, Yanchao Yang, Yi Ma
FoCLIP is a framework that creates adversarial examples to manipulate CLIP-based image quality metrics by reducing the alignment between image and text features. It uses stochastic gradient descent to combine feature alignment, score distribution balancing, and pixel‑guard regularization, enabling high CLIPscore predictions while maintaining visual fidelity. Experiments on artistic prompts and ImageNet show significant CLIPscore gains, and the authors also propose a color‑channel sensitivity detection method that achieves 91% accuracy.
By Yulin Chen, Zeyuan Wang, Tianyuan Yu, Yingmei Wei, Liang Bai
Efficient Multi-Modal Planning with Reward-Guided Preference Optimization for Autonomous Driving proposes EMPlan, a hybrid trajectory planning method that combines sparse anchors with an offset refinement module for low-latency, high-accuracy predictions. The approach uses a two-stage training paradigm—pretraining followed by reward-guided fine-tuning—to improve safety without extra inference cost, leveraging rule-based reward signals and unpaired preference supervision. EMPlan is evaluated on the non-reactive NAVSIM benchmark, achieving a favorable balance between planning accuracy and efficiency under real-time constraints.
By Chenglin Chen, Lujia Wang, Xinhu Zheng, Jun Ma, Haoang Li
DiffWAM is a geometry‑conditioned navigation world‑action model that transforms predictive features from a frozen video foundation model into continuous camera trajectories, eliminating the need for future‑video synthesis and multi‑frame reconstruction during deployment. Its Grid‑Motion module preserves spatial‑temporal motion associations, while Latent2Pose grounds them with first‑frame geometry to recover metrically meaningful 3D motion. The system, complemented by FastDreamer for asynchronous trajectory handoff, achieves a trajectory RMSE of 0.3492 m and a 74.40 % endpoint success rate on the DiffWAM‑1000 benchmark, with real‑world tests showing complex UAV behaviors and an onboard implementation reaching 1.08 s latency on NVIDIA Jetson AGX Thor.
By Mo Zhu, Yuze Wu, Xijie Huang, Xiao Cui, Fei Gao, Xin Zhou
arXiv:2609.38864v1 Announce Type: new
Abstract: Embodied tasks demand accurate, flexible, and semantically rich 3D scene representations. 3D semantic occupancy is well suited to this requirement, as...
By Jinglong Wang, Yunjie Wang, Zhiyang Zhang, Jiawei He, Ye Yuan, Bo Qiu, Jing Zhang
arXiv:2609.39623v1 Announce Type: new
Abstract: The proliferation of high-fidelity generative editing models has made it possible to inject violent or sexual content into otherwise ordinary images wh...
By Yoonseo Kim, Seungwoo Baek, Junyoung Park
arXiv:2609.38285v1 Announce Type: cross
Abstract: Vision-language models (VLMs) can contradict themselves across views of the same spatial relation and fail to respond when that relation changes. Add...
By Hongbo Wang, Zihan Lin, Wenkui Yang, Shiran Ge, Yuang Ai, Jie Cao, Huaibo Huang, Ran He
arXiv:2609.39145v1 Announce Type: cross
Abstract: Unreliable visual inputs can harm task performance and cause potential physical safety risks for vision-language-action (VLA) models. We analyze how...
By Heejae Suh, Jongwook Han, Zahra Gholami, Yohan Jo
arXiv:2609.39304v1 Announce Type: cross
Abstract: When an off-the-shelf coding agent is used directly as a robot policy, observing a browser-based 3D interface through screenshots and acting by posin...
By Zhijie Wei, Ferris Tan, Jinghui Wang
arXiv:2609.39599v1 Announce Type: cross
Abstract: 3D visuomotor policies provide a strong foundation for spatially precise manipulation, yet current text-to-3D policies struggle to follow unseen fine...
By Xinhao Yang, Wenhao Wu, Ning Lv, Yanshen Ding, Zhenhong Sun, Daoyi Dong, Chunlin Chen, Zhi Wang
arXiv:2609.38620v1 Announce Type: new
Abstract: Neural implicit representations have had a significant impact on scene reconstruction by enabling robots to build continuous, differentiable, and high-...
By Hanwen Cao, Wenqiang Wu, Kuang-Ting Tu, Mathias Otnes, Jeffrey Delmerico, Rui Wang, Yulun Tian, Nikolay Atanasov