arXiv AI By Geoffroy Keime, Nicolas Cuperlier, Benoit R. Cottereau

REACT: A Fully Spiking State-Space Model for Real-Time Event-Driven Temporal Perception

Read the original on arXiv AI →

REACT is a fully spiking state‑space model that processes raw event‑camera data one event at a time, avoiding temporal accumulation and its associated delay. It employs a complex‑valued spiking neuron (C‑SiLIF) whose dynamics are driven by the inter‑event interval, enabling continuous‑time state updates at microsecond resolution. Evaluated on gesture recognition and time‑to‑collision estimation, REACT achieves low latency (4.6 ms) and high accuracy, supports anytime prediction, zero‑shot transfer, and INT8 quantization, dramatically reducing energy consumption.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
3d ago

Spike-driven Vision-Language-Action Model

arXiv:2609.39514v1 Announce Type: new Abstract: Vision-language-action (VLA) models bridge multimodal understanding and robotic control, advancing the dominant paradigm for embodied intelligence. How...

By Shuai Wang, Malu Zhang, Mingquan Liu, Weihui Dai, Dehao Zhang, Jieyuan Zhang, Yimeng Shan, Zijian Zhou, Yang Yang
arXiv Computer Vision
Sep 24

Bend the Clock: Predicting Ahead to Beat Latency in Event-Based Object Detection

The paper introduces ChronoFuse, a causal availability-time detector that predicts object states at the time its output becomes available rather than at the observation timestamp, addressing the latency mismatch in event-based multi-object detection. ChronoFuse performs lightweight cross-time fusion over a multi-scale feature hierarchy, adding only 0.17 M parameters and 0.84 ms latency overhead. It recovers a large portion of accuracy lost to latency, achieving up to 20.95 sAP on EV‑Flying data compared to 2.25 sAP for the strongest standard detector.

By Biswadeep Sen, Benoit R. Cottereau, Nicolas Cuperlier, Terence Sim
arXiv Machine Learning
Aug 27

LM-X: Explainable Action Modeling with Progress, Event, and Uncertainty Prediction for Generalist Robot Manipulation

LM‑X is a generalist vision‑language‑action policy that augments action prediction with three online, explicitly supervised signals: return‑to‑go (RTG) for task progress, event‑to‑go (ETG) for the next semantic transition, and heteroscedastic action flow for local reliability. By conditioning action generation on these signals, LM‑X embeds explainability directly into control rather than as a post‑hoc explanation. After a 20‑day pretraining run on 64 GPUs, LM‑X outperforms an action‑only backbone by 16.0 points and a single‑head variant by 10.8 points, and achieves 74.1 % success on 50 RoboTwin2.0 tasks and 68.6 % on seven real‑robot tasks, surpassing the GR00T N1.7 baseline.

By Jin Lou, Jingxuan Zhu, Andong Chen, Xupeng Wang, Yuan Xu, Yuexuan Li, Xingdong Zhu, Zhijie Zhu, Yingwei Ji, Wenpeng Nie, Jingyi Li, Liangliang Chen, Jinyan Liu, Zhiqi Song, Jidong Zhang, Hongming Li, Yuchen Zhu