arXiv:2610.01283v1 Announce Type: new
Abstract: Reliable shelf monitoring is an important capability for retail automation, yet existing out-of-stock detection methods mainly operate in image space a...
By Lingyi Zhou, Yunke Wang, Mengyu Zheng, Wenbo Wang, Zijian Wang, Chang Xu
arXiv:2610.01544v1 Announce Type: new
Abstract: Existing image forgery detectors often suffer from generalization to unseen manipulation methods due to the limited ability to capture transferable for...
By Bingjian Yang, Shilei Zhao, Zheng Wang
arXiv:2610.01741v1 Announce Type: new
Abstract: Predictive Vision-Language-Action (VLA) models aim to improve robotic manipulation via future observation or world dynamics forecasting. However, exist...
By Yijie Zhu, Rui Shao, Jie He, Wei Li, Bo Zhao, Yelin Wang, Xiaochen Yuan, Tao Tan, Miao Zhang, Xiaojiang Peng, Zitong Yu
arXiv:2610.01778v1 Announce Type: new
Abstract: Reliable evaluation of image forgery localization (IFL) requires assessing models under diverse distribution changes, yet existing benchmarks often cov...
By Baoke Dou, Ziye Wang, Hao Wang, Guoqing Cai, Wende Tan, Chenyang Si, Liucheng Guo, Yueming Lyu
arXiv:2610.01939v1 Announce Type: new
Abstract: Vision language model (VLM) agents can control robots through visual feedback and action primitives, but repeated model invocations and redundant obser...
By Ruiyang Si, Jianxin Bi, Shunyu Yang, Rui Ni, Wenbo Huang, Qiang Wang, Shulong Jiang, Duomin Wang, Xiuyu Li, Haiwen Feng, Zhen Dong, Daquan Zhou
arXiv:2610.00360v1 Announce Type: cross
Abstract: Reinforcement learning (RL) for dexterous manipulation must discover finger-object contacts and then control the object precisely; the action noise t...
By Haoyu Wang, Siyuan Qian, Yanjun Li, Zeyu Zhang, Yandong Guo, Boxin Shi, Hao Tang
arXiv:2610.00330v1 Announce Type: cross
Abstract: Long-term object search requires learning where objects usually appear from repeated but uneven observations of a changing environment. We formulate...
By Jiaming Wang, Zhiwei Xue, Chen Jizhuo, Peng Shiqi, Harold Soh
arXiv:2610.00981v1 Announce Type: cross
Abstract: We focus on language-conditioned flow-based manipulation, where robot flows (robot velocity fields) serve as embodiment-agnostic, motion-centric repr...
By Shota Kobayashi, Koki Seno, Daichi Yashima, Komei Sugiura
arXiv:2610.01742v1 Announce Type: cross
Abstract: Equipping artificial agents with spatial intelligence requires a comprehensive generative prior over the dynamic 3D world. We propose World Motion Mo...
By Jiahui Lei, Qianqian Wang, Trevor Darrell, Angjoo Kanazawa
arXiv:2610.02196v1 Announce Type: cross
Abstract: We study test-time evolution for humanoid loco-manipulation: solving tasks that a controller was never trained for by repurposing its existing skills...
By Zhuo Lin, Sirui Xu, Liuyu Bian, Yu-Xiong Wang, Liang-Yan Gui
arXiv:2606.21562v2 Announce Type: replace
Abstract: Transformers are AI's workhorse but their computational cost becomes prohibitive when processing long sequences. We target long-horizon streaming v...
By Philippe Weinzaepfel, Christian Wolf, Mert B\"ulent Sariyildiz, Guillaume Bono, Gianluca Monaci
The paper reviews end‑to‑end autonomous driving (E2E‑AD) training, framing it as a Data‑Strategy‑Platform system. It surveys recent advances in data pipelines, learning paradigms, and training infrastructures, and discusses how these layers interact to influence model performance, robustness, and deployability. The authors highlight current limitations and propose a future vision that prioritizes data value, foundation‑driven generalization, and integrated training‑testing loops for more robust, scalable, and trustworthy autonomous driving systems.
By Chengkai Xu, Yiming Cui, Jiaqi Liu, Yicheng Guo, Cheng Qin, Geyuan Zhang, Xinwei Dong, Shiyu Fang, Peng Hang, Jian Sun
The paper introduces GTA‑VLA, an interactive Vision‑Language‑Action framework that lets users guide robot policies with explicit visual cues such as affordance points, boxes, and traces. Unlike traditional direct sense‑to‑act models, GTA‑VLA incorporates a spatial‑visual Chain‑of‑Thought that blends human guidance with internal task planning, and couples this reasoning module with a lightweight reactive action head for efficient execution. Experiments on the SimplerEnv WidowX benchmark show a state‑of‑the‑art 81.2 % success rate, and the framework significantly improves task success under out‑of‑domain visual shifts and spatial ambiguities, demonstrating the benefit of interactive reasoning for failure recovery in embodied control.
By Yiran Ling, Qing Lian, Jinghang Li, Qing Jiang, Tianming Zhang, Xiaoke Jiang, Chuanxiu Liu, Jie Liu, Lei Zhang
The paper introduces an action‑conditioned Network World Model that learns how a network’s diffusion dynamics evolve under interventions over time. This model can quickly predict the outcomes of actions, enabling a coding agent to design and refine algorithms that select actions to maximize expected performance on complex network tasks. Experiments on eight network tasks and five diffusion models show that the resulting algorithms match or surpass the best existing baselines in 138 of 141 settings while achieving up to 14.5× faster rollouts than traditional Monte Carlo simulation.
By Rishab Alagharu, Hongji Pu, Zeeshan Memon, Xinyuan Song, Yuntong Hu, Liang Zhao
The paper introduces a bounded‑fidelity sim‑as‑demo‑stage design pattern that suppresses contact physics during object handoffs in simulators, using MuJoCo’s mocap‑body primitive and a lightweight Python adapter. This approach ensures audit‑chain stability, producing identical event‑log hashes across 1,000 replays per posture, whereas a contact‑force baseline yields significant divergence. The authors demonstrate that the pattern maintains reproducibility across various timesteps and sequential handoffs with minimal overhead, and they identify contexts where it should not be applied.
By Xue Qin, Simin Luan, Cong Yang, Zhijun Li
The paper introduces a method for generating legible plans in arbitrary PDDL domains by extending prior legibility research to classical planning without custom planners. It incorporates a second‑order theory of mind to estimate the observer’s perspective, enabling robots to implicitly communicate goals in human‑robot teaming. Benchmark results show that increasing legibility typically trades off with plan efficiency, and a regularizing factor is needed to balance the two.
By Michele Persiani, Thomas Hellstr\"om
arXiv:2610.00604v1 Announce Type: cross
Abstract: Vision-language-action policies often see only one or a few recent frames, which makes it difficult to evaluate how they use information that disappe...
By Egor Cherepanov, Nikita Kachaev, Aleksandr I. Panov, Alexey K. Kovalev
arXiv:2610.00899v1 Announce Type: cross
Abstract: Autoregressive Vision-Language-Action models often represent continuous robot actions as discrete token sequences, enabling action prediction with st...
By Keisuke Shirai, Tomohiro Motoda, Hanbit Oh, Ryoichi Nakajo, Roman Mykhailyshyn, Ryo Hanai, Shotaro Miwa, Yukiyasu Domae
arXiv:2610.01260v1 Announce Type: cross
Abstract: Quadrupedal locomotion requires balancing conflicting objectives such as command tracking, stability, and energy efficiency, yet conventional reinfor...
By Amr Mousa, Rifny Rachman, Neil Karavis, Michele Caprio, Richard Allmendinger
arXiv:2610.01559v1 Announce Type: cross
Abstract: World Action Models (WAMs) predict visual futures and robot actions, yet they remain susceptible to task-incomplete imagination, where plausible, act...
By Seungyeon Kim, Junhoo Lee, Baekseung Kim, Minkyu Kim, Nojun Kwak