Hugging Face Trending Papers

FLEET: Token-Based Feature Extraction for Event Camera-based Reinforcement Learning

Event cameras generate asynchronous, high-frequency data streams offering spatially sparse information at lower latency than traditional cameras. In principle, these properties should be ideal for the design of control policies.

arXiv Computer Vision
Aug 27

FLEET: Token-Based Feature Extraction for Event Camera-based Reinforcement Learning

FLEET is a token‑based feature extractor that processes event camera data directly, using random Fourier features and cross‑attention to compress variable‑length event streams into fixed‑size latent representations. By decoupling inference cost from sensor resolution, it avoids the high compute and temporal blurring associated with CNN‑based grid aggregation. Experiments on a new high‑throughput benchmark show that FLEET outperforms state‑of‑the‑art methods and remains robust across different observation frequencies.

By Tristan Gottwald, Maximilian Schier, Melanie Schaller, Bodo Rosenhahn
arXiv Computer Vision
Aug 27

Low-Latency Event-Based Object Detection with Spatially-Sparse Linear Attention

The paper introduces Spatially‑Sparse Linear Attention (SSLA), a novel attention mechanism that activates only a sparse subset of spatial states, enabling efficient parallel training and inference for event‑based vision. Building on SSLA, the authors present SSLA‑Det, an end‑to‑end asynchronous linear attention model that achieves state‑of‑the‑art accuracy on Gen1 and N‑Caltech101 while reducing per‑event computation by more than 20× compared to the strongest prior asynchronous baseline.

By Haiqing Hao, Zhipeng Sui, Rong Zou, Zijia Dai, Nikola Zubi\'c, Davide Scaramuzza, Wenhui Wang
arXiv Machine Learning
Sep 7

LookThere! Sparse Vision by Reinforced Selection

LookThere! Sparse Vision by Reinforced Selection proposes an end‑to‑end reinforcement learning framework that jointly trains a shallow input selector and a deep representation extractor for vision transformers. The selector learns where to focus and the extractor learns what to process, enabling the model to use only a tiny fraction of the input tokens—down to 0.2%—while preserving accuracy. The method outperforms existing selection techniques across diverse tasks and models, including high‑resolution recognition, segmentation, zero‑shot classification, and regression, establishing a new Pareto frontier in performance‑compute trade‑offs.

By Sreehari Rammohan, Yousef Yassin, Anthony Fuller, Junfeng Wen, Carl Vondrick, Evan Shelhamer
arXiv AI
Sep 18

PointEvent: Rethinking Event-based Tiny Object Detection via Serialized Motion Evidence Accumulation

PointEvent introduces serialized motion evidence accumulation for event-based tiny object detection, treating motion continuity as an ordered evidence propagation process. The method organizes event streams into locality‑preserving spatiotemporal paths and chronology‑preserving temporal paths, alternating serialized scans across complementary orders to consolidate fragmented motion evidence. A lightweight event‑wise state‑space framework with a high‑resolution event branch and compact context modulation achieves state‑of‑the‑art performance with the fewest parameters and fastest inference among compared methods.

By Zongze Wu, Baofeng Jia, Weiqi Yan, Jingyuan Zhang, Yu Zang, Xiaoyu Chen, Jing Han
arXiv Computer Vision
Sep 7

Efficient Multi-Timescale Event Representations for Feed-Forward Object Detection

The paper introduces a confidence‑normalized continuous multi‑timescale representation for event cameras, using logarithmic B‑spline temporal encoding and a geometry‑aware local confidence mechanism. When paired with a fixed feed‑forward EventCenterNet detector, this representation outperforms the compact CSTR representation on the PEDRo and Gen1 datasets. Additionally, a recursive exponential‑polynomial approximation is proposed to allow efficient event‑by‑event updates while maintaining detection performance.

By Fredrik Lundell, Per-Erik Forssen, M{\aa}rten Wadenb\"ack, Astrid Lundmark
arXiv Computation and Language
Sep 3

ShallowStream: Index Shallow then Answer Deep for Streaming Video Understanding

ShallowStream is a framework for streaming video understanding that uses the shallow layers of a multimodal large language model (MLLM) to encode frames and build a lightweight index. During streaming, it maintains an always‑on index via the KV cache of shallow layers, and at query time it scores context frames using shallow‑layer attention and selects diverse evidence for answering. The approach matches the performance of leading streaming methods while cutting per‑frame prefill latency and 10‑second end‑to‑end latency by up to 52.1× and 11.9×, respectively.

By Jitai Hao, Ke Yang, Qiang Huang, Jun Yu
arXiv Machine Learning
Sep 22

VISTA: An Attention-Based Multi-Agent Reinforcement Learning Architecture for Space Situational Awareness Sensor Tasking

arXiv:2609.23875v1 Announce Type: new Abstract: The rapid growth of resident space objects is increasing the complexity of space situational awareness sensor tasking, challenging classical optimizati...

By Miguel Leiva-V\'elez, Adalberto Claudio Quiros, Nicolas Gaston Rozado, Hodei Urrutxua, V\'ictor Rodr\'iguez-Fern\'andez
arXiv Computer Vision
Aug 28

Parameter Efficient Continual Learning for Sparse Event-Based Transformers

The paper introduces sLoTh, a parameter‑efficient continual learning framework for sparse event‑based vision transformers. sLoTh freezes the backbone and limits plasticity to low‑rank attention updates (seLoRA) and shared neuronal threshold modulation, updating less than 1% of parameters without replay buffers. Experiments on CIFAR‑100, Tiny‑ImageNet, ImageNet‑100, and ImageNet‑R show competitive rehearsal‑free performance across up to 100 tasks while achieving roughly 6.5× lower energy consumption than dense vision transformers.

By Vaishnavi Nagabhushana, Kartikay Agrawal, Ayon Borthakur