arXiv:2606. 31421v1 Announce Type: cross Abstract: Single-stage video object detectors are increasingly deployed in time-critical applications, yet it remains unclear whether these models genuinely reason over temporal context or merely exploit a single informative frame-a gap hidden by standard metrics, which reward correct predictions regardless of how they are reached.
By Karam Tomotaki-Dawoud, Anna Hilsmann, Peter Eisert, Sebastian Bosse
REACT is a fully spiking state‑space model that processes raw event‑camera data one event at a time, avoiding temporal accumulation and its associated delay. It employs a complex‑valued spiking neuron (C‑SiLIF) whose dynamics are driven by the inter‑event interval, enabling continuous‑time state updates at microsecond resolution. Evaluated on gesture recognition and time‑to‑collision estimation, REACT achieves low latency (4.6 ms) and high accuracy, supports anytime prediction, zero‑shot transfer, and INT8 quantization, dramatically reducing energy consumption.
By Geoffroy Keime, Nicolas Cuperlier, Benoit R. Cottereau
arXiv:2609.16864v1 Announce Type: cross
Abstract: Vision-language-action (VLA) models have achieved impressive performance in quasi-static manipulation, but struggle in dynamic manipulation tasks bec...
By Zhenyang Feng, Jimin Heo, Erik B. Sudderth, Unnat Jain
arXiv:2512. 01031v2 Announce Type: replace-cross Abstract: Vision-Language-Action models (VLAs) are becoming increasingly capable across diverse robotic tasks.
By Jiaming Tang, Yufei Sun, Yilong Zhao, Shang Yang, Yujun Lin, Zhuoyang Zhang, James Hou, Yao Lu, Zhijian Liu, Song Han
arXiv:2605. 21862v2 Announce Type: replace-cross Abstract: Chunked vision-language-action (VLA) policies predict multi-step robot controls, conditioning each update on the current visual observation alone.
By Chushan Zhang, Ruihan Lu, Jinguang Tong, Xuesong Li, Yikai Wang, Hongdong Li
arXiv:2605. 19294v2 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) policies increasingly rely on asynchronous inference to hide large-model latency behind ongoing robot motion.
By Yixiang Zhu, Yonghao Chen, Zijie Yang, Yusong Hu, Xinyu Chen