arXiv Machine Learning

A2SG:Adaptive and Asymmetric Surrogate Gradients for Training Deep Spiking Neural Networks

arXiv:2606. 11236v1 Announce Type: cross Abstract: Training deep spiking neural networks (SNNs) remains challenging due to sharp loss landscapes and temporal inconsistency caused by surrogate gradients.

arXiv AI
Aug 17

SAGE: Surrogate-gradient Adaptation via Attention-Guided Entropy for Spiking Transformers

arXiv:2608. 13702v1 Announce Type: cross Abstract: Spiking neural networks (SNNs) offer an energy-efficient alternative to conventional deep neural networks by exploiting sparse event-driven computation, but their training remains challenging because the non-differentiable spike function requires surrogate gradients whose fixed shape may be suboptimal across layers and training stages.

By Kiran Nair, Rodrigue Rizk, KC Santosh
arXiv AI
Jun 19

Hybrid ANN-SNN Pipeline with Local Plasticity

arXiv:2606. 20151v1 Announce Type: cross Abstract: This work proposes a hybrid ANN-SNN pipeline that effectively leverages the rich embeddings of pretrained artificial neural networks (ANNs) to enable high-performance spiking neural networks (SNNs).

By Denis Larionov, Khairutin Shtanchaev, Mikhail Kiselev, Mikhail Korovin, Ivan Tugoy
arXiv Machine Learning
Sep 25

On the second-order optimization for spiking neural networks

The paper introduces SpiKFAX, a second‑order optimization technique for Spiking Neural Networks (SNNs) that uses a Kronecker‑factored approximation of the Fisher information matrix tailored to the sparse, discrete, and temporally recurrent dynamics of SNNs. By addressing the sharp loss landscape that hampers training with conventional optimizers, SpiKFAX improves test accuracy and training stability across five architectures and seven datasets. The method offers a computationally tractable alternative to existing curvature‑based approaches for SNNs.

By Ngoc Phu Doan, Ihsen Alouani
arXiv AI
2d ago

Contrastive Attention Mitigates Spectral Bias in Spiking Transformers

The paper introduces Spiking Contrastive Attention (SCA), a module designed to reduce spectral bias in Spiking Transformers by enhancing high‑frequency information. It demonstrates that spiking neurons and spiking self‑attention act as low‑pass filters, leading to loss of high‑frequency components. Experiments show that SCA improves performance across image classification, semantic segmentation, and event‑based tracking while maintaining lower complexity than the original spiking self‑attention.

By Xiaoli Liu, Malu Zhang, Yang Yang
arXiv Machine Learning
Aug 19

Spikformer V2: Join the High Accuracy Club on ImageNet with an SNN Ticket

Spikformer V2 introduces a Spiking Self‑Attention (SSA) mechanism that removes softmax and uses spike‑based Query, Key, and Value to capture sparse visual features efficiently. It also adds a Spiking Convolutional Stem (SCS) and employs self‑supervised learning (masking and reconstruction) to pre‑train the model before fine‑tuning on ImageNet. The result is the first spiking neural network to surpass 80 % accuracy on ImageNet, achieving 81.10 % with a 172 M‑parameter, 16‑layer model in just one time step.

By Zhaokun Zhou, Yijie Lu, Kaiwei Che, Wei Fang, Keyu Tian, Qihao Peng, Yuesheng Zhu, Shuicheng Yan, Yonghong Tian, Li Yuan
arXiv AI
Jul 15

Burst Spiking Neural Networks

arXiv:2607. 11914v1 Announce Type: cross Abstract: A central goal of current Spiking Neural Network (SNN) research is to improve their accuracy toward becoming low-power alternatives to Artificial Neural Networks (ANNs).

By Jiahong Zhang, Sijun Shen, Man Yao, Han Xu, Mingqiang Huang, Yonghong Tian, Bo Xu, Guoqi Li
arXiv Machine Learning
Sep 15

Exploring napping paradigm for Recurrent Spiking Neural Networks

The paper proposes a biologically inspired micro‑sleep technique called napping for recurrent spiking neural networks, combining proportional weight scaling with continuous stochastic membrane activity. Experiments on an unsupervised SNN trained with trace‑based STDP on Gabor‑preprocessed MNIST show that well‑tuned napping can match the classification accuracy of conventional weight normalization while offering different clustering characteristics. The study suggests that napping may be preferable when representational structure is more important than raw classification speed, despite its higher simulation cost.

By Andreas Massey, Stefano Nichele, Aliaksandr Hubin