arXiv AI

A Neuromorphic Trigger for Efficient Audio Event Detection

arXiv:2606. 17775v1 Announce Type: cross Abstract: Efficient processing of continuous audio streams remains a key challenge for real-time and resource-constrained systems.

arXiv AI
Jun 24

Neuromorphic Speech Enhancement with Dual-Branch Spiking Neural Networks

arXiv:2606. 23761v1 Announce Type: cross Abstract: Spiking neural network (SNN)-based neuromorphic speech enhancement has emerged as a promising paradigm due to its energy efficiency, yet it still underperforms classical artificial neural network (ANN)-based approaches owing to binary activations and the lack of well-designed network architectures.

By Taiyu Meng, Wenbin Jiang, Haoyi Zhang, Yuhan Zhou, Haibing Yin
arXiv AI
Aug 20

Low-Power, Neuromorphic, Acoustic Anomaly Detection for Persistent Machine Monitoring

The paper presents an autoencoder-based acoustic anomaly detection system implemented on Intel’s Loihi 2 neuromorphic processor. It achieves high detection performance—0.9959 AUC on a clean ToyADMOS ToyCar benchmark and 0.7990 source AUC on a noisy DCASE 2026 Task 2 ToyCar benchmark—while operating with only 0.0406–0.0426 mJ of dynamic energy per sample, far below CPU and GPU baselines. The system demonstrates that low‑power, on‑chip inference is feasible for persistent machine monitoring.

By Steven C. Nesbit (Information Sciences, CAI-3, Los Alamos National Laboratory, Los Alamos, USA), Victor M. Vergara (AeroVironment Inc., Albuquerque, USA), Michael A. Felix (University of New Mexico COSMIAC Research Center, Albuquerque, USA), Evan T. Kain (Air Force Research Laboratory, Kirtland AFB, USA), Luis R. Garc\'ia Carrillo (Air Force Research Laboratory, Kirtland AFB, USA), Gerd J. Kunde (Nuclear and Particle Physics and Applications, P-3, Los Alamos National Laboratory, Los Alamos, USA), Andrew T. Sornborger (Information Sciences, CAI-3, Los Alamos National Laboratory, Los Alamos, USA)
arXiv AI
Sep 12

Investigating catastrophic forgetting in sound event classification

The paper explores methods to mitigate catastrophic forgetting in incremental learning for sound event classification. It evaluates architectural and regularization strategies on FSD50K and AudioSet, finding that deeper layers, especially the classifier head, are most vulnerable. The most effective approach identified is fully freezing the feature extractor while fine‑tuning a dynamic head, which achieves minimal forgetting, stable training, and a balanced trade‑off between memory stability and learning plasticity.

By Riccardo Casciotti, Annamaria Mesaros
arXiv AI
Jul 7

Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding

arXiv:2607. 04383v1 Announce Type: cross Abstract: Large Audio-Language Models (LALMs) reason fluently about sound yet struggle to localize precisely when events occur, while classical Sound Event Detection attains frame-level precision only over a closed label set.

By Zihan Zhang, Xize Cheng, Wenhao Yan, Tong Zhang, Dongjie Fu, Boyun Zhang, Yongbo He, Tao Jin
arXiv Computation and Language
Sep 7

Large Language Models with At Most One Spike per Neuron

The paper presents a spiking neural network (SNN) approach that uses time-to-first-spike (TTFS) coding to limit each neuron to at most one spike per time window, enabling energy-efficient large language models (LLMs). A reference-based strategy is introduced to encode the four core LLM components—embedding layers, layer normalization, attention-related operations, and dropout—allowing the construction of a fully TTFS-based SNN architecture trained end-to-end. Experiments on BERT and GPT-2 show performance comparable to artificial neural network (ANN) counterparts on natural language understanding and common-sense reasoning, while achieving a 1.5‑billion‑parameter spiking LLM and providing an estimate of spike-related energy consumption.

By Zhuoya Zhao, Parsa Omidi, Aref Jafari, Richard Naud
arXiv Machine Learning
Sep 22

NAVIR: Neuromorphic Audio-Visual Speech Recognition for Robust Human-Robot Interaction on Edge Hardware

NAVIR is an end‑to‑end audio‑visual speech recognition system designed for the BrainChip Akida neuromorphic processor, which only supports sequential 2‑D convolutions. The architecture separates spatial and temporal encoding into three AkidaNet modules—per‑frame visual, temporal video, and spectrogram audio encoders—fused by a lightweight predictor and decoded with constrained beam search. Trained with CTC on noise‑augmented audio and fine‑tuned via quantization‑aware training, the quantized model achieves 14.0% WER on GRID’s unseen‑speaker split and 3.3% on overlapped‑speaker split, outperforming audio‑only baselines, and delivers 98.6% command accuracy at 1.5% WER on an industrial‑command corpus, while offering a 13‑fold energy advantage over conventional ANNs and roughly 5‑fold lower energy per inference than a Raspberry Pi CPU.

By Leonidas Delimpasis, Panagiota Moraiti, Antonis Porichis, Panos Chatzakos, Michail Karamousadakis
arXiv AI
Jun 2

DAStatFormer: A Hybrid Multibranch Transformer with Statistical Feature Integration for DAS-Based Pattern Recognitions

arXiv:2606. 00081v1 Announce Type: cross Abstract: Distributed Acoustic Sensing (DAS) enables large-scale monitoring through optical fibers, but its high dimensionality and complex spatio-temporal patterns make event classification demanding.

By Michel Dione (CERI SN - IMT Nord Europe), Jerry Lonlac (CERI SN - IMT Nord Europe), H\'el\`ene Louis (CERI SN - IMT Nord Europe), Anthony Fleury (CERI SN - IMT Nord Europe), Stephane Lecoeuche