The paper presents a redesign of the LiSenNet speech‑enhancement model for deployment on the STM32N6570‑DK Neural‑ART microcontroller accelerator. By replacing the recurrent bottleneck with convolutional mixers, converting unsupported operations to static int8 primitives, and using bounded decoder activations, the authors achieve an NPU‑compatible model that matches or surpasses the original LiSenNet in quality (PESQ 3.08 vs 3.01 FP32) while running each 16 ms input hop in 4.83 ms (real‑time factor 0.30). The study demonstrates that co‑designing parameter count, operator compatibility, quantization range, and streaming state is essential for efficient real‑time speech enhancement on constrained NPUs.
By Cl\'ement Laroche, Rasmus Kongsgaard Olsson
arXiv:2608.30792v1 Announce Type: cross
Abstract: Obtaining data from neuromorphic sensors and processing it with Spiking Neural Networks is a promising solution to lower the energy cost of artificia...
By Valentin M. Meunier, Am\'elie Gruel, Pierre Lewden, Adrien F. Vincent, Sylvain Sa\"ighi
The paper introduces the Neural Action Codec (NAC), a convolutional encoder‑decoder architecture that treats short robot action trajectories as multi‑channel 1D signals and compresses them using a multi‑scale residual vector quantization (RVQGAN) model. NAC replaces traditional discrete action tokenizers with a compact, ordered token space via offset codebooks, allowing standard autoregressive policies to operate over short, structured sequences while a Vocos‑style decoder reconstructs the actions. Experiments on LIBERO‑10, RoboMimic, and real‑world manipulation tasks show that NAC achieves higher reconstruction fidelity and better average success rates than existing binning, FAST, and VQ‑based tokenizers at comparable or improved compression rates.
By Ahad Jawaid, Yu Xiang
The paper introduces Training-Free Omni (TFO), a plug‑and‑play framework that transforms a frozen vision‑language model (VLM) into a speech‑centric omni model without modifying its architecture or requiring multimodal re‑alignment. TFO leverages Whisper to generate confidence‑filtered, timestamped transcripts and routes them through the VLM’s existing language interface, leaving the visual pathway untouched. Evaluations on 56 benchmarks across 21 languages show that TFO matches or surpasses native omni models on audio‑visual tasks, improves audio‑only performance, and preserves strong visual and reasoning capabilities.
By Ankan Deria, Hanoona Rasheed, Xilin He, Fahad Shahbaz Khan, Salman Khan
arXiv:2603.16086v2 Announce Type: replace-cross
Abstract: While recent Vision-Language-Action (VLA) models have begun to incorporate audio, they typically treat sound as static pre-execution prompts...
By Chang Nie, Tianchen Deng, Guangming Wang, Zhe Liu, Hesheng Wang
arXiv:2606. 12688v1 Announce Type: cross Abstract: We are entering a new era of composite model architectures that integrate diverse components such as vision encoders, language backbones, diffusion and flow heads, audio codecs, action generators, and world-model predictors.
By Atindra Jha, Naomi Sagan, Keisuke Kamahori, Irmak Sivgin, Rohan Sanda, Steven Gao, Mark Horowitz, Luke Zettlemoyer, Olivia Hsu, Jure Leskovec, Baris Kasikci, Stephanie Wang