arXiv Machine Learning

NAVIR: Neuromorphic Audio-Visual Speech Recognition for Robust Human-Robot Interaction on Edge Hardware

NAVIR is an end‑to‑end audio‑visual speech recognition system designed for the BrainChip Akida neuromorphic processor, which only supports sequential 2‑D convolutions. The architecture separates spatial and temporal encoding into three AkidaNet modules—per‑frame visual, temporal video, and spectrogram audio encoders—fused by a lightweight predictor and decoded with constrained beam search. Trained with CTC on noise‑augmented audio and fine‑tuned via quantization‑aware training, the quantized model achieves 14.0% WER on GRID’s unseen‑speaker split and 3.3% on overlapped‑speaker split, outperforming audio‑only baselines, and delivers 98.6% command accuracy at 1.5% WER on an industrial‑command corpus, while offering a 13‑fold energy advantage over conventional ANNs and roughly 5‑fold lower energy per inference than a Raspberry Pi CPU.

arXiv Machine Learning
Sep 25

Beyond Model Size: Redesigning LiSenNet for embedded speech enhancement

The paper presents a redesign of the LiSenNet speech‑enhancement model for deployment on the STM32N6570‑DK Neural‑ART microcontroller accelerator. By replacing the recurrent bottleneck with convolutional mixers, converting unsupported operations to static int8 primitives, and using bounded decoder activations, the authors achieve an NPU‑compatible model that matches or surpasses the original LiSenNet in quality (PESQ 3.08 vs 3.01 FP32) while running each 16 ms input hop in 4.83 ms (real‑time factor 0.30). The study demonstrates that co‑designing parameter count, operator compatibility, quantization range, and streaming state is essential for efficient real‑time speech enhancement on constrained NPUs.

By Cl\'ement Laroche, Rasmus Kongsgaard Olsson
arXiv Machine Learning
5d ago

NAC: Neural Action Codec for Vision-Language-Action Models

The paper introduces the Neural Action Codec (NAC), a convolutional encoder‑decoder architecture that treats short robot action trajectories as multi‑channel 1D signals and compresses them using a multi‑scale residual vector quantization (RVQGAN) model. NAC replaces traditional discrete action tokenizers with a compact, ordered token space via offset codebooks, allowing standard autoregressive policies to operate over short, structured sequences while a Vocos‑style decoder reconstructs the actions. Experiments on LIBERO‑10, RoboMimic, and real‑world manipulation tasks show that NAC achieves higher reconstruction fidelity and better average success rates than existing binning, FAST, and VQ‑based tokenizers at comparable or improved compression rates.

By Ahad Jawaid, Yu Xiang
arXiv Computer Vision
Sep 7

Training-Free Speech-Centric Omni Understanding with Frozen VLMs

The paper introduces Training-Free Omni (TFO), a plug‑and‑play framework that transforms a frozen vision‑language model (VLM) into a speech‑centric omni model without modifying its architecture or requiring multimodal re‑alignment. TFO leverages Whisper to generate confidence‑filtered, timestamped transcripts and routes them through the VLM’s existing language interface, leaving the visual pathway untouched. Evaluations on 56 benchmarks across 21 languages show that TFO matches or surpasses native omni models on audio‑visual tasks, improves audio‑only performance, and preserves strong visual and reasoning capabilities.

By Ankan Deria, Hanoona Rasheed, Xilin He, Fahad Shahbaz Khan, Salman Khan
arXiv AI
Jun 12

M*: A Modular, Extensible, Serving System for Multimodal Models

arXiv:2606. 12688v1 Announce Type: cross Abstract: We are entering a new era of composite model architectures that integrate diverse components such as vision encoders, language backbones, diffusion and flow heads, audio codecs, action generators, and world-model predictors.

By Atindra Jha, Naomi Sagan, Keisuke Kamahori, Irmak Sivgin, Rohan Sanda, Steven Gao, Mark Horowitz, Luke Zettlemoyer, Olivia Hsu, Jure Leskovec, Baris Kasikci, Stephanie Wang
arXiv Machine Learning
Sep 22

AVTR-1: Open Stack for Real-Time Interactive Avatars

arXiv:2609.22913v1 Announce Type: cross Abstract: Talking-head and dyadic models now achieve real-time inference, yet fast motion generation alone does not produce an interactive conversation. A live...

By Artem Kravtsov, Dmitrii Ziganshin, Vsevolod Poletaev, Gleb Balitskiy, Anastasia Tikhonova, Egor Burkov, Vadim Lebedev
arXiv Computer Vision
Sep 1

DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution

arXiv:2608.31106v1 Announce Type: new Abstract: Recent video generators often omit audio or synthesize it in a separate stage, limiting reciprocal modeling of visual dynamics and acoustic events. We...

By Jiashu Zhu, Yanhao Zheng, Ruitian Tian, Rujing Dang, Shen Zhang, Bingze Song, Jiachen Lei, Ruimin Lin, Jiahong Wu, Xiangxiang Chu
arXiv AI
Sep 17

GrainSpeech: Less Context, More Detail for Compact Speech Synthesis

GrainSpeech is a compact speech synthesis model that uses a fixed‑receptive‑field convolutional encoder to reduce pitch, energy, and duration prediction errors by 36.0%, 17.3%, and 3.4% respectively. It introduces a Mel‑specific gradient‑variance supervision that improves fine‑scale variation while avoiding quality degradation. With only 264.8K parameters, GrainSpeech achieves 17.9× real‑time Mel generation on a microcontroller and attains UTMOS scores comparable to much larger models, using less than 1.5% of their parameters.

By Zitao Liang, Chang Gao