arXiv:2604. 08894v2 Announce Type: replace-cross Abstract: Spiking Neural Networks (SNNs) offer superior energy efficiency over Artificial Neural Networks (ANNs).
By Zecheng Hao, Shenghao Xie, Kang Chen, Wenxuan Liu, Zhaofei Yu, Tiejun Huang
arXiv:2608. 19238v1 Announce Type: cross Abstract: Spiking Transformers model token interactions primarily through spiking self-attention (SSA).
By Dongcheng Zhao, Sicheng Shen, Zhenyu Yang, Zhiyuan Li, Jinyan Yu, Yongjian Wang, Tiechui Yao, Wenli Zhang, Tielin Zhang
arXiv:2605. 08270v2 Announce Type: replace-cross Abstract: Spiking Neural Networks (SNNs) offer notable advantages in biological plausibility and energy efficiency, making them promising candidates for building low-power Transformers.
By Zequan Xie, Weiming Zeng, Yunhua Chen, Sichang Ling, Tongyang Chen, Jinsheng Xiao
Spikformer V2 introduces a Spiking Self‑Attention (SSA) mechanism that removes softmax and uses spike‑based Query, Key, and Value to capture sparse visual features efficiently. It also adds a Spiking Convolutional Stem (SCS) and employs self‑supervised learning (masking and reconstruction) to pre‑train the model before fine‑tuning on ImageNet. The result is the first spiking neural network to surpass 80 % accuracy on ImageNet, achieving 81.10 % with a 172 M‑parameter, 16‑layer model in just one time step.
By Zhaokun Zhou, Yijie Lu, Kaiwei Che, Wei Fang, Keyu Tian, Qihao Peng, Yuesheng Zhu, Shuicheng Yan, Yonghong Tian, Li Yuan
arXiv:2606. 12287v1 Announce Type: cross Abstract: The Transformer architecture is widely regarded as the most powerful tool for natural language processing, but due to a high number of complex operations, it inherently faces the issue of high energy consumption.
By Claas Beger, Florian Walter, Alois Knoll
arXiv:2608. 13702v1 Announce Type: cross Abstract: Spiking neural networks (SNNs) offer an energy-efficient alternative to conventional deep neural networks by exploiting sparse event-driven computation, but their training remains challenging because the non-differentiable spike function requires surrogate gradients whose fixed shape may be suboptimal across layers and training stages.
By Kiran Nair, Rodrigue Rizk, KC Santosh