Spike-Aware C++ INT8 Inference for Sparse Spiking Language Models on Commodity CPUs
arXiv:2606. 03026v1 Announce Type: cross Abstract: Spiking language models expose activation sparsity that dense Transformer runtimes do not directly exploit.
The paper introduces an implicit-perturbation zeroth-order (IPZO) architecture for fine-tuning spiking transformers on in‑memory computing (IMC) accelerators. By generating perturbations only for spike‑activated weight rows and combining them with IMC weighted sums, the design eliminates costly read‑modify‑write operations and reduces the hardware footprint of random number generators. An address‑driven XOR recombination scheme (PGU‑XOR) further mitigates spatial correlations, achieving near‑software accuracy while cutting perturbation energy by up to 50% compared to conventional explicit perturbation methods.
arXiv:2606. 03026v1 Announce Type: cross Abstract: Spiking language models expose activation sparsity that dense Transformer runtimes do not directly exploit.
arXiv:2606. 30676v1 Announce Type: cross Abstract: Deploying spiking neural networks (SNNs) on neuromorphic hardware demands aggressive synaptic pruning while preserving temporal computation integrity.
arXiv:2604. 08894v2 Announce Type: replace-cross Abstract: Spiking Neural Networks (SNNs) offer superior energy efficiency over Artificial Neural Networks (ANNs).
arXiv:2606.03026v2 Announce Type: replace-cross Abstract: Binary spike activations allow a language-model runtime to read only active weight columns and replace multiplications by weight sums. We imp...
arXiv:2607. 14672v1 Announce Type: new Abstract: Continuous-time spiking neural networks (SNNs) provide an event-driven framework for temporal computation, computational neuroscience, and neuromorphic hardware.
arXiv:2605.30361v2 Announce Type: replace-cross Abstract: Spiking Neural Networks (SNNs) offer compelling energy efficiency on neuromorphic hardware, yet their training remains challenging because th...
arXiv:2608. 08479v1 Announce Type: cross Abstract: Spiking neural networks (SNNs) offer a promising pathway to energy-efficient AI and brain-inspired computing.
arXiv:2606. 17249v1 Announce Type: cross Abstract: The dominant trajectory of modern machine learning has been to scale up: larger models, larger accelerators, larger memory budgets.
arXiv:2601. 22876v2 Announce Type: replace Abstract: Spiking neural networks (SNNs) promise energy-efficient inference for large language models (LLMs), yet most reported savings rely on compute-operation counts that overlook data movement.
arXiv:2606.06159v2 Announce Type: replace-cross Abstract: Spiking neural networks (SNNs) have the potential to emerge as the third generation of neural networks and have attracted increasing attentio...
arXiv:2409. 08290v5 Announce Type: replace-cross Abstract: Spiking Neural Networks (SNNs) promise higher energy efficiency over conventional Quantized Artificial Neural Networks (QNNs) due to their event-driven, spike-based computation.
arXiv:2606. 06159v1 Announce Type: cross Abstract: Spiking neural networks (SNNs) have the potential to emerge as the third generation of neural networks and have attracted increasing attention across a wide range of applications.