Spiking Local Interaction and Adaptive Complementary Fusion for Spiking Transformer
arXiv:2608. 19238v1 Announce Type: cross Abstract: Spiking Transformers model token interactions primarily through spiking self-attention (SSA).
arXiv:2608. 19238v1 Announce Type: cross Abstract: Spiking Transformers model token interactions primarily through spiking self-attention (SSA).
Infrared and visible image fusion (IVIF) integrates the complementary information of two modalities into a single image with richer scene content. While existing methods are largely built on artificial neural networks (ANNs), which densely compute over all activations, spiking neural networks (SNNs) communicate through sparse binary spikes and compute only where and when a spike occurs, offering a route to more energy-efficient fusion.
arXiv:2606. 31135v1 Announce Type: cross Abstract: We present LINet (Linear Integration Network), a Multi-Stream Neural Network (MSNN) for RGB-D scene classification.
Spikformer V2 introduces a Spiking Self‑Attention (SSA) mechanism that removes softmax and uses spike‑based Query, Key, and Value to capture sparse visual features efficiently. It also adds a Spiking Convolutional Stem (SCS) and employs self‑supervised learning (masking and reconstruction) to pre‑train the model before fine‑tuning on ImageNet. The result is the first spiking neural network to surpass 80 % accuracy on ImageNet, achieving 81.10 % with a 172 M‑parameter, 16‑layer model in just one time step.
arXiv:2607. 14672v1 Announce Type: new Abstract: Continuous-time spiking neural networks (SNNs) provide an event-driven framework for temporal computation, computational neuroscience, and neuromorphic hardware.
Visual Prompting (VP) has emerged as an efficient paradigm for adapting large-scale pre-trained vision models to downstream tasks by incorporating learnable prompts at the input level. However, existing VP methods typically employ dense pixel-level prompts, which often suffer from redundant perturbations, limited generalization and energy inefficiency.
arXiv:2605. 05895v2 Announce Type: replace-cross Abstract: Modern AI-generated videos are photorealistic at the single-frame level, leaving inter-frame dynamics as the main remaining axis for detection.
arXiv:2606. 20151v1 Announce Type: cross Abstract: This work proposes a hybrid ANN-SNN pipeline that effectively leverages the rich embeddings of pretrained artificial neural networks (ANNs) to enable high-performance spiking neural networks (SNNs).
arXiv:2609.39514v1 Announce Type: new Abstract: Vision-language-action (VLA) models bridge multimodal understanding and robotic control, advancing the dominant paradigm for embodied intelligence. How...
SBMVTrack is a fully spiking neural network framework designed for energy-efficient UAV visual tracking. It introduces Energy-Weighted Spike Budgeting (EWSB) to constrain spike activity based on computational cost, and Masked Multi-View Target Modeling (MVTM) to enhance target representation by leveraging correlated temporal views. Experiments on multiple benchmarks show that SBMVTrack reduces spike firing rates and theoretical energy consumption while maintaining competitive tracking accuracy.
arXiv:2607.01630v2 Announce Type: replace Abstract: Dynamic expansion methods for class-incremental learning (CIL) protect task-specific knowledge by growing dedicated tokens or subnetworks, yet our...
arXiv:2608. 03324v1 Announce Type: new Abstract: Federated learning enables collaborative model training across distributed edge devices while strictly preserving data privacy.