arXiv:2609.25146v1 Announce Type: new
Abstract: Continual learning, the ability to learn from sequential experience while retaining and adapting prior knowledge, is central to intelligent systems ope...
By Hongwei Yan, Kanglei Zhou, Qi Cheng, Weiyi Dong, Chunyan Lan, Guanglong Sun, Jun Zhou, Qian Li, Yi Zhong, Liyuan Wang
arXiv:2609.25836v1 Announce Type: new
Abstract: Multi-task optimization (MTO) addresses a set of optimization tasks simultaneously, often suffering from inaccurate inter-task relationship estimation...
By Tingyang Wei, Haofeng Wu, Jiao Liu, Zhao Wei, Puay Siew Tan, Yew-Soon Ong
arXiv:2609.26231v1 Announce Type: new
Abstract: Continuous and subtle GNSS spoofing poses a serious threat to autonomous vehicles because forged positions may remain locally plausible while gradually...
By Muhammad Ayub Sabir, Junbiao Pang, Fatima Ashraf
arXiv:2609.25654v1 Announce Type: cross
Abstract: Robots operating safely in cluttered everyday environments often need to infer scene geometry from partial observations. Methods that detect objects...
By Dongwon Son, Junhyek Han, Yoontae Cho, Minseok Lee, Hong-seok Choi, Jiwook Choi, Hyungjin Kim, Beomjoon Kim
arXiv:2609.25017v1 Announce Type: new
Abstract: Deepfakes, synthetic audiovisual content produced by deep generative models, have escalated into a critical threat across civilian and military domains...
By Alexandros Gazis, Efstathios Karypidis, Kleanthi Santamouri, Theodoros Vavouras, Nikos E. Mastorakis, Stylianos Pappas
arXiv:2609.25627v1 Announce Type: cross
Abstract: General-purpose robot control requires models to understand task intent, identify where to interact, capture how the scene evolves, and generate prec...
By Haoran Wen, Wenfu Wang, Kunsong Shi, Jingke Wang, Wancheng Feng, Yiren Zhang, Yueran Zhao, Xuancheng Zhang, Nanfei Ye, Xingru Chen, Zhaohong Sun, Chengmin Yang, Zikang Yu, Penghao Bi, Jia Shi, Yu Liu, Kun Zhan, Yan Xie
arXiv:2609.25831v1 Announce Type: cross
Abstract: Recent VLM-based autonomous driving planners adopt GRPO-style reinforcement learning to optimize driving performance. However, existing GRPO recipes...
By Yuqi Ye, Shangkun Sun, Junhong Lin, Jiayi Zhao, Changhao Peng, Wei Zheng, Guoqing Liu, Tiesong Zhao, Wei Gao
arXiv:2512.02697v4 Announce Type: replace
Abstract: Cross-view geo-localization infers a location by retrieving geo-tagged reference images matching a query image. However, the traditional satellite-...
By Zixuan Song, Jing Zhang, Di Wang, Zhiming Luo, Wenbin Liu, Haonan Guo, En Wang, Bo Du, Liangpei Zhang
arXiv:2605.25059v4 Announce Type: replace
Abstract: Crucial for autonomous exploration, online 3D occupancy prediction and mapping incrementally construct dense spatial representations on the fly. Em...
By Ruoyu Wang, Yong Liu, Jiahan Li, Sheng Tao, Yuhang Lin, Yukai Ma
arXiv:2609.24526v2 Announce Type: replace
Abstract: Physical AI requires models to ground visual and linguistic understanding in real-world environments while accounting for environmental constraints...
By Foundation Model, Li Auto Inc
arXiv:2607.04745v2 Announce Type: replace-cross
Abstract: Strong mapped-region thermal visual place recognition (VPR) does not ensure safe rejection of unmapped queries. We identify and quantify this...
By Zhiyuan Lu, Kanji Tanaka
DUMA-Bench is a new benchmark that evaluates the security of large language model agents in dual‑control settings, where both the agent and the user can modify the shared environment. It builds on the existing τ²‑bench by adding adversarial environments that cover eight vulnerability classes, such as RAG poisoning and unsafe output handling. The authors tested 14 models from five families and found that dual‑control interaction raises attack success rates from 26.9% to 41.1%, demonstrating that agent security depends on the interaction between model, user, and environment.
By Ivan Aleksandrov, German Kochnev, Sabrina Sadiekh, Yaroslav Rogoza
RiverVLN introduces the first benchmark for long‑horizon vision‑language navigation (VLN) of unmanned surface vehicles (USVs) in continuous riverine motion. The PGT‑NAV framework converts navigation instructions into an ordered sequence of visually verifiable semantic phases, maintaining an active phase online through grounded visual and motion evidence. This phase‑grounded approach reduces recursive position and heading drift, achieving a 0.79 success rate in Unity‑ROS closed‑loop tests and demonstrating transfer to real‑world USV deployment.
By Jieling Wu, Yuehao Huang, Jiajun Lv, Tao Huang, Yong Liu, Weiwei Liu
The study evaluates whether large language models (LLMs) used as synthetic personas can predict real audience responses to marketing copy. Using thousands of headline A/B tests from the Upworthy Research Archive, the authors compare a ten-persona panel grounded in real audience demographics to a no-persona zero‑shot baseline that asks the model for a typical reader’s click likelihood. Results show that the no‑persona baseline outperforms the persona‑based approach, with higher predictive validity and top‑1 accuracy, indicating that forcing the model to role‑play specific personas introduces bias and noise.
By Alexandre Cristov\~ao Maiorano
MatchFusion is a learnable module that performs explicit-implicit instance matching for spatio‑temporal multimodal autonomous driving. It initializes pairwise affinities with geometric similarity and category consistency, then refines associations using instance embeddings to guide a residual aggregation operator for adaptive information exchange. Experiments on nuScenes show that MatchFusion improves perception accuracy, reduces FLOPs by 55.3% and GPU memory usage by 39.3%, and adds only 3.7% of total perception latency.
By Xiaoyu Li, Jiajia Fu, Long Shi, Tianyu Du, Ruihang Li, Xian Wu, Lijun Zhao, Yingtao Zhang, Lining Sun, Ruifeng Li
The paper presents a spiking neural network-based Proximal Policy Optimization (SNN‑PPO) algorithm for autonomous UAV navigation through constrained openings in civil infrastructure. By integrating spike‑based actor‑critic reinforcement learning with a stochastic Gaussian policy, the method achieved 63.77 % overall success across 3000+ episodes, improving to over 90 % in later stages and averaging 2.10 windows per episode.
By Francis Noah Walugembe, Maciej Wielgosz, Toma\v{z} Gori\v{c}an, Matej Mertik
HybridFlow is a generative policy for robotic manipulation that uses a three‑stage inference procedure requiring only two network function evaluations (2‑NFE). The policy first generates a coarse action trajectory with a Global Jump based on MeanFlow, then refines the state using a parameter‑free ReNoise interpolation, and finally performs a Local Refine to query the instantaneous‑velocity limit. Experiments on RoboMimic and five real‑robot settings show that HybridFlow achieves high success rates and improves task performance over a 16‑step Diffusion Policy while reducing action‑generation latency by roughly eightfold.
By Zhenchen Dong, Fulin Chen, Jinna Fu, Jiaming Wu, Qingran Wu, Shengyuan Yu, Hongyu Yu, Yide Liu
SAIL is a framework that transforms robot imitation learning into an iterative refinement problem, enabling test-time scaling of trajectory generation. It employs Monte Carlo Tree Search where each node represents a full trajectory and edges denote refinements, guided by an archive of successful trajectories, a vision‑language model for scoring, and step‑level feedback. Experiments on six manipulation tasks in simulation and real‑world settings show that higher test‑time compute consistently raises success rates, reaching up to 95% on complex tasks.
By Makoto Sato, Yusuke Iwasawa, Yujin Tang, So Kuroki
Metric-Bench introduces a new benchmark for Vision‑Language Models (VLMs) that focuses on metric‑spatial reasoning in indoor scenes by using in‑image reference objects with known dimensions. The accompanying MetricReasoner fine‑tuning recipe employs structured prompts and numerical rewards to implicitly learn 2D‑to‑3D mapping without camera intrinsics. Experiments show that this approach improves spatial metric understanding by 43.1 % over existing models and boosts downstream embodied tasks, while also delivering gains on general VLM benchmarks.
By Yuling Xi, Haokai Zhang, Muzhi Zhu, Hao Zhong, Zongze Du, Hengyu Zhao, Chenchen Jing, Yufei Yin, Bin Qin, Yongjie Yang, Zhenbo Luo, Hao Chen, Chunhua Shen
The paper "End-to-End Visual Odometry with RNNs and Attention" presents a study of deep‑learning approaches to visual odometry (VO), proposing a novel temporal attention‑based model to enhance performance. It evaluates existing end‑to‑end VO methods and explores their effectiveness on hand‑held camera data, contrasting with the typical driving‑scene training sets. The work aims to improve VO accuracy in more dynamic and complex visual environments.
By Ruiyu Li, Yinjia Liu, Alexander Yu