arXiv:2608. 19238v1 Announce Type: cross Abstract: Spiking Transformers model token interactions primarily through spiking self-attention (SSA).
By Dongcheng Zhao, Sicheng Shen, Zhenyu Yang, Zhiyuan Li, Jinyan Yu, Yongjian Wang, Tiechui Yao, Wenli Zhang, Tielin Zhang
arXiv:2609.37047v1 Announce Type: cross
Abstract: We propose a new spiking neural network (SNN) design to process static images and event streams using time-to-first-spike (TTFS) latencies. Our key r...
By Aidin Attar, Eleonora Cicciarella, Michele Rossi
Spikformer V2 introduces a Spiking Self‑Attention (SSA) mechanism that removes softmax and uses spike‑based Query, Key, and Value to capture sparse visual features efficiently. It also adds a Spiking Convolutional Stem (SCS) and employs self‑supervised learning (masking and reconstruction) to pre‑train the model before fine‑tuning on ImageNet. The result is the first spiking neural network to surpass 80 % accuracy on ImageNet, achieving 81.10 % with a 172 M‑parameter, 16‑layer model in just one time step.
By Zhaokun Zhou, Yijie Lu, Kaiwei Che, Wei Fang, Keyu Tian, Qihao Peng, Yuesheng Zhu, Shuicheng Yan, Yonghong Tian, Li Yuan
The paper introduces Spiking Contrastive Attention (SCA), a module designed to reduce spectral bias in Spiking Transformers by enhancing high‑frequency information. It demonstrates that spiking neurons and spiking self‑attention act as low‑pass filters, leading to loss of high‑frequency components. Experiments show that SCA improves performance across image classification, semantic segmentation, and event‑based tracking while maintaining lower complexity than the original spiking self‑attention.
By Xiaoli Liu, Malu Zhang, Yang Yang
arXiv:2609.39514v1 Announce Type: new
Abstract: Vision-language-action (VLA) models bridge multimodal understanding and robotic control, advancing the dominant paradigm for embodied intelligence. How...
By Shuai Wang, Malu Zhang, Mingquan Liu, Weihui Dai, Dehao Zhang, Jieyuan Zhang, Yimeng Shan, Zijian Zhou, Yang Yang
arXiv:2606. 31135v1 Announce Type: cross Abstract: We present LINet (Linear Integration Network), a Multi-Stream Neural Network (MSNN) for RGB-D scene classification.
By Gabriel Clinger