Spiking Local Interaction and Adaptive Complementary Fusion for Spiking Transformer
arXiv:2608. 19238v1 Announce Type: cross Abstract: Spiking Transformers model token interactions primarily through spiking self-attention (SSA).
Infrared and visible image fusion (IVIF) integrates the complementary information of two modalities into a single image with richer scene content. While existing methods are largely built on artificial neural networks (ANNs), which densely compute over all activations, spiking neural networks (SNNs) communicate through sparse binary spikes and compute only where and when a spike occurs, offering a route to more energy-efficient fusion.
arXiv:2608. 19238v1 Announce Type: cross Abstract: Spiking Transformers model token interactions primarily through spiking self-attention (SSA).
arXiv:2609.37047v1 Announce Type: cross Abstract: We propose a new spiking neural network (SNN) design to process static images and event streams using time-to-first-spike (TTFS) latencies. Our key r...
Spikformer V2 introduces a Spiking Self‑Attention (SSA) mechanism that removes softmax and uses spike‑based Query, Key, and Value to capture sparse visual features efficiently. It also adds a Spiking Convolutional Stem (SCS) and employs self‑supervised learning (masking and reconstruction) to pre‑train the model before fine‑tuning on ImageNet. The result is the first spiking neural network to surpass 80 % accuracy on ImageNet, achieving 81.10 % with a 172 M‑parameter, 16‑layer model in just one time step.
The paper introduces Spiking Contrastive Attention (SCA), a module designed to reduce spectral bias in Spiking Transformers by enhancing high‑frequency information. It demonstrates that spiking neurons and spiking self‑attention act as low‑pass filters, leading to loss of high‑frequency components. Experiments show that SCA improves performance across image classification, semantic segmentation, and event‑based tracking while maintaining lower complexity than the original spiking self‑attention.
arXiv:2609.39514v1 Announce Type: new Abstract: Vision-language-action (VLA) models bridge multimodal understanding and robotic control, advancing the dominant paradigm for embodied intelligence. How...
arXiv:2606. 31135v1 Announce Type: cross Abstract: We present LINet (Linear Integration Network), a Multi-Stream Neural Network (MSNN) for RGB-D scene classification.
arXiv:2606. 20151v1 Announce Type: cross Abstract: This work proposes a hybrid ANN-SNN pipeline that effectively leverages the rich embeddings of pretrained artificial neural networks (ANNs) to enable high-performance spiking neural networks (SNNs).
arXiv:2607. 11914v1 Announce Type: cross Abstract: A central goal of current Spiking Neural Network (SNN) research is to improve their accuracy toward becoming low-power alternatives to Artificial Neural Networks (ANNs).
SpikeMoE introduces a spike-based k‑WTA router that uses lateral inhibition and refractory periods to select the top‑K experts based on discrete spike counts, inspired by hippocampal CA1 competition. The framework combines spiking neural network dynamics with mixture‑of‑experts conditional computation and adds a two‑stage missing‑modality module for robust multimodal processing. Experiments on vision, language, and multimodal tasks show that SpikeMoE matches or surpasses ANN baselines while offering energy‑efficient performance.
arXiv:2604. 08894v2 Announce Type: replace-cross Abstract: Spiking Neural Networks (SNNs) offer superior energy efficiency over Artificial Neural Networks (ANNs).
arXiv:2606. 11236v1 Announce Type: cross Abstract: Training deep spiking neural networks (SNNs) remains challenging due to sharp loss landscapes and temporal inconsistency caused by surrogate gradients.
arXiv:2606. 13016v1 Announce Type: new Abstract: Spiking neural networks (SNNs) are promising for energy-efficient inference, and time-to-first-spike (TTFS) coding is especially attractive because each neuron fires at most once.