arXiv:2609.24403v1 Announce Type: new
Abstract: Biological visual systems achieve continuous, low-latency motion perception by processing sparse, asynchronous spiking signals, enabling real-time trac...
By Mazdak Fatahi, \v{S}\'arka Pryjmakov\'a, Pierre Boulet, Giulia D'Angelo
arXiv:2609.39514v1 Announce Type: new
Abstract: Vision-language-action (VLA) models bridge multimodal understanding and robotic control, advancing the dominant paradigm for embodied intelligence. How...
By Shuai Wang, Malu Zhang, Mingquan Liu, Weihui Dai, Dehao Zhang, Jieyuan Zhang, Yimeng Shan, Zijian Zhou, Yang Yang
arXiv:2510.26614v2 Announce Type: replace
Abstract: We propose tokenization of events and present a tokenizer, Spiking Patches, specifically designed for event cameras. Given a stream of asynchronous...
By Christoffer Koo {\O}hrstr{\o}m, Ronja G\"uldenring, Lazaros Nalpantidis
The paper introduces ChronoFuse, a causal availability-time detector that predicts object states at the time its output becomes available rather than at the observation timestamp, addressing the latency mismatch in event-based multi-object detection. ChronoFuse performs lightweight cross-time fusion over a multi-scale feature hierarchy, adding only 0.17 M parameters and 0.84 ms latency overhead. It recovers a large portion of accuracy lost to latency, achieving up to 20.95 sAP on EV‑Flying data compared to 2.25 sAP for the strongest standard detector.
By Biswadeep Sen, Benoit R. Cottereau, Nicolas Cuperlier, Terence Sim
arXiv:2608. 11605v1 Announce Type: new Abstract: World Action Models (WAMs) couple future visual prediction with robot action generation, enabling policies to model how the physical world evolves during interaction.
By Jiakai Huang, Zhongbo Wu, Zheng Zhang, Zihan Wang, Shan You, Tao Huang
LM‑X is a generalist vision‑language‑action policy that augments action prediction with three online, explicitly supervised signals: return‑to‑go (RTG) for task progress, event‑to‑go (ETG) for the next semantic transition, and heteroscedastic action flow for local reliability. By conditioning action generation on these signals, LM‑X embeds explainability directly into control rather than as a post‑hoc explanation. After a 20‑day pretraining run on 64 GPUs, LM‑X outperforms an action‑only backbone by 16.0 points and a single‑head variant by 10.8 points, and achieves 74.1 % success on 50 RoboTwin2.0 tasks and 68.6 % on seven real‑robot tasks, surpassing the GR00T N1.7 baseline.
By Jin Lou, Jingxuan Zhu, Andong Chen, Xupeng Wang, Yuan Xu, Yuexuan Li, Xingdong Zhu, Zhijie Zhu, Yingwei Ji, Wenpeng Nie, Jingyi Li, Liangliang Chen, Jinyan Liu, Zhiqi Song, Jidong Zhang, Hongming Li, Yuchen Zhu
arXiv:2606.20092v3 Announce Type: replace
Abstract: Memory remains a critical bottleneck for long-horizon robotic manipulation, as standard Vision-Language-Action (VLA) policies often fail when task-...
By Ganlin Yang, Zhangzheng Tu, Yuqiang Yang, Sitong Mao, Junyi Dong, Tianxing Chen, Jiaqi Peng, Jing Xiong, Jiafei Cao, Jifeng Dai, Wengang Zhou, Yao Mu, Tai Wang
arXiv:2512. 01031v2 Announce Type: replace-cross Abstract: Vision-Language-Action models (VLAs) are becoming increasingly capable across diverse robotic tasks.
By Jiaming Tang, Yufei Sun, Yilong Zhao, Shang Yang, Yujun Lin, Zhuoyang Zhang, James Hou, Yao Lu, Zhijian Liu, Song Han
arXiv:2606. 31167v1 Announce Type: cross Abstract: VLA models have emerged as a powerful paradigm for transferring semantic knowledge from web-scale data to physical robotic control.
By Hao Sun, Yu Song, Shiyu Teng, Ziwei Niu, Yen-Wei Chen
arXiv:2609.16864v1 Announce Type: cross
Abstract: Vision-language-action (VLA) models have achieved impressive performance in quasi-static manipulation, but struggle in dynamic manipulation tasks bec...
By Zhenyang Feng, Jimin Heo, Erik B. Sudderth, Unnat Jain
StreamPI introduces a streaming multimodal temporal modeling framework that enhances Vision‑Language‑Action models by adding temporal reasoning without extra parameters. It anchors each visual observation and language instruction pair as a temporal unit, using bidirectional attention for cross‑modal fusion and causal attention for autoregressive streaming inference. The method employs random‑interval streaming training to improve robustness and leverages the LLM backbone’s length extrapolation to inherit pretrained weights, achieving superior performance over pi0.5 on real‑robot and simulation tasks.
By Zhe Liu, Jinghua Hou, Yuxiang Lu, Zhenya Yang, Xianzhe Fan, Junwei Luo, Junyi Li, Ruihua Han, Zhi Hou, Hengshuang Zhao
The paper introduces sLoTh, a parameter‑efficient continual learning framework for sparse event‑based vision transformers. sLoTh freezes the backbone and limits plasticity to low‑rank attention updates (seLoRA) and shared neuronal threshold modulation, updating less than 1% of parameters without replay buffers. Experiments on CIFAR‑100, Tiny‑ImageNet, ImageNet‑100, and ImageNet‑R show competitive rehearsal‑free performance across up to 100 tasks while achieving roughly 6.5× lower energy consumption than dense vision transformers.
By Vaishnavi Nagabhushana, Kartikay Agrawal, Ayon Borthakur