arXiv AI

Neuromorphic Speech Enhancement with Dual-Branch Spiking Neural Networks

arXiv:2606. 23761v1 Announce Type: cross Abstract: Spiking neural network (SNN)-based neuromorphic speech enhancement has emerged as a promising paradigm due to its energy efficiency, yet it still underperforms classical artificial neural network (ANN)-based approaches owing to binary activations and the lack of well-designed network architectures.

arXiv Machine Learning
Jun 5

DBHN-Net: Dual-Branch Hybrid Neural Network For Low-Complexity Monaural Speech Enhancement

arXiv:2606. 05911v1 Announce Type: cross Abstract: Although artificial neural network (ANN) based speech enhancement (SE) methods demonstrate excellent performance, the high computational complexity and high energy consumption hinder their deployment in practical front-end processing tasks.

By Cunhang Fan, Enrui Liu, Jing Zhou, Jian Kang, Jie Li, Andong Li, Jian Zhou, Zhao Lv, Xuelong Li
arXiv AI
Jun 24

End-to-End Radar and Communication Modulation Recognition with Neuromorphic Computing

arXiv:2606. 24075v1 Announce Type: cross Abstract: Although deep learning-based methods can achieve high accuracy in automatic modulation recognition (AMR) tasks, their high computational cost makes it difficult to strike a balance between accuracy and power consumption, thereby limiting their application on resource-constrained platforms.

By Xiaohu Li, Chongxiao Qu, Caiyong Lin, Chenxiao Dou, Wei Hua
arXiv Machine Learning
Sep 25

On the second-order optimization for spiking neural networks

The paper introduces SpiKFAX, a second‑order optimization technique for Spiking Neural Networks (SNNs) that uses a Kronecker‑factored approximation of the Fisher information matrix tailored to the sparse, discrete, and temporally recurrent dynamics of SNNs. By addressing the sharp loss landscape that hampers training with conventional optimizers, SpiKFAX improves test accuracy and training stability across five architectures and seven datasets. The method offers a computationally tractable alternative to existing curvature‑based approaches for SNNs.

By Ngoc Phu Doan, Ihsen Alouani
arXiv AI
2d ago

Contrastive Attention Mitigates Spectral Bias in Spiking Transformers

The paper introduces Spiking Contrastive Attention (SCA), a module designed to reduce spectral bias in Spiking Transformers by enhancing high‑frequency information. It demonstrates that spiking neurons and spiking self‑attention act as low‑pass filters, leading to loss of high‑frequency components. Experiments show that SCA improves performance across image classification, semantic segmentation, and event‑based tracking while maintaining lower complexity than the original spiking self‑attention.

By Xiaoli Liu, Malu Zhang, Yang Yang
arXiv Machine Learning
Sep 22

NAVIR: Neuromorphic Audio-Visual Speech Recognition for Robust Human-Robot Interaction on Edge Hardware

NAVIR is an end‑to‑end audio‑visual speech recognition system designed for the BrainChip Akida neuromorphic processor, which only supports sequential 2‑D convolutions. The architecture separates spatial and temporal encoding into three AkidaNet modules—per‑frame visual, temporal video, and spectrogram audio encoders—fused by a lightweight predictor and decoded with constrained beam search. Trained with CTC on noise‑augmented audio and fine‑tuned via quantization‑aware training, the quantized model achieves 14.0% WER on GRID’s unseen‑speaker split and 3.3% on overlapped‑speaker split, outperforming audio‑only baselines, and delivers 98.6% command accuracy at 1.5% WER on an industrial‑command corpus, while offering a 13‑fold energy advantage over conventional ANNs and roughly 5‑fold lower energy per inference than a Raspberry Pi CPU.

By Leonidas Delimpasis, Panagiota Moraiti, Antonis Porichis, Panos Chatzakos, Michail Karamousadakis
arXiv Computation and Language
Sep 7

Large Language Models with At Most One Spike per Neuron

The paper presents a spiking neural network (SNN) approach that uses time-to-first-spike (TTFS) coding to limit each neuron to at most one spike per time window, enabling energy-efficient large language models (LLMs). A reference-based strategy is introduced to encode the four core LLM components—embedding layers, layer normalization, attention-related operations, and dropout—allowing the construction of a fully TTFS-based SNN architecture trained end-to-end. Experiments on BERT and GPT-2 show performance comparable to artificial neural network (ANN) counterparts on natural language understanding and common-sense reasoning, while achieving a 1.5‑billion‑parameter spiking LLM and providing an estimate of spike-related energy consumption.

By Zhuoya Zhao, Parsa Omidi, Aref Jafari, Richard Naud