arXiv AI By Sunday Afariogun, Odunolaoluwa Jenrola, Zeinab Nezami

Offline AI Modules: Voice-First Offline Architecture, Hardware Reference Stack, Quantization and Benchmarking

Read the original on arXiv AI →

The paper introduces the Offline AI Modules workstream, which provides a voice‑first offline architecture, a low‑cost hardware reference bill of materials, and a reproducible quantization and benchmarking pipeline for instruction‑tuned language models in the 2‑5B parameter range. It evaluates three models across four quantization formats on two hardware tiers—NVIDIA Jetson Orin NX and Raspberry Pi5—measuring deployment metrics and multilingual quality. The key result is that Q4_K_M quantization offers the best size‑to‑quality trade‑off, enabling high decode throughput and strong topic classification accuracy while staying within memory limits on both tiers.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 29

LUMO (Lightweight Unified Multilingual Orchestrator): A Privacy Preserving Offline Voice Assistant

LUMO (Lightweight Unified Multilingual Orchestrator) is a privacy‑preserving offline voice assistant that runs entirely on edge hardware, specifically a Raspberry Pi 5 with 8 GB RAM. It integrates local ASR, a 4‑bit GGUF‑quantized LLM, and TTS to deliver end‑to‑end response latencies of 2.0–4.0 s, a 6.8 % WER on short English utterances, and lower peak power consumption (~9 W) compared to existing edge assistants. The system also supports Bangla speech, enabling multilingual use in low‑resource settings.

By Md. Mehedi Hasan Naeem, Mst. Kamrunnahar Ruma, Nafiza Anjum, Shakila Sultana, Md. Sujan Ali
arXiv Machine Learning
Sep 25

Same Bit Width, Different Outcomes: Post-Training Quantization of Text-to-Speech Across Architectures

The paper evaluates post‑training quantization (PTQ) for text‑to‑speech (TTS) models across multiple architectures using a unified protocol. It shows that reducing weights to 4‑bit per‑channel can significantly lower predicted mean opinion scores (UTMOS) and that even 8‑bit per‑tensor scaling can cause severe degradation, with the impact varying by model. A staged ablation identifies the sensitive components, and per‑layer GPTQ can recover performance to within 0.1 UTMOS, while real int8 and int4 kernels confirm the simulated results on hardware, demonstrating that each configuration must be validated on the target runtime.

By Se Un Park, Yutae Kim, Junyoung Park
arXiv Machine Learning
Sep 10

TontaubeV1: Streaming Text-to-Speech with Hierarchical Codec Modeling and Bounded Context

TontaubeV1 is a streaming text‑to‑speech model that preserves natural prosody while running on a single consumer GPU. It encodes speech with a hierarchical DualCodec representation at 12.5 Hz, separating a semantic stream from successive acoustic refinements, and uses Qwen3‑derived transformers to predict the semantic stream, utterance duration, and acoustic refinements. The system supports up to one minute of reference audio for voice conditioning, streams with a 200 ms latency to first audio, and achieves real‑time factors of 0.08 (single input) and 0.02 (eight concurrent inputs).

By Fritz Cremer, Jonathan Cremer
Hugging Face Trending Papers
Sep 8

TontaubeV1: Streaming Text-to-Speech with Hierarchical Codec Modeling and Bounded Context

TontaubeV1 is a streaming text‑to‑speech model that maintains natural prosody while running on a single consumer GPU. It encodes speech with a hierarchical DualCodec representation at 12.5 Hz, separating a semantic stream from successive acoustic refinements, and uses Qwen3‑derived transformers to predict the semantic stream and add refinements. The model supports up to one minute of reference audio for voice conditioning, streams with a 200 ms latency to first audio, and achieves real‑time factors of 0.08 (single input) and 0.02 (eight concurrent inputs).

arXiv AI
Sep 21

NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities

arXiv:2609.21967v1 Announce Type: cross Abstract: We introduce NemotronLabs VoiceChat, an open full-duplex speech-to-speech model with native tool-calling capabilities. NemotronLabs VoiceChat combine...

By Jagadeesh Balam, Travis Bartley, Edresson Casanova, Sanjay Chauhan, Chen Chen, Zhehuai Chen, Zijia Chen, Francesco Ciannella, Slyne Deng, Mikyas Desta, Harishchandra Dubey, Slim Essid, Nourchene Ferchichi, Boris Ginsburg, Mariana Graterol Fuenmayor, Negar Habibi, Kevin Hu, Anand Joseph, Viraj Karandikar, Myungjong Kim, Viacheslav Klimkov, Seelan Lakshmi Narasimhan, Lily Lee, Jason Li, Eileen Long, Ameya Mahabaleshwarkar, Aditya Malte, Adi Margolin, Sasha Meister, Valentin Mendelev, Oluwatobi Olabiyi, Ankita Pasad, Yifan Peng, Elena Rastorgueva, Jayda Ritchie, Jason Roche, Nikhil Srihari, Yuanhang Su, Yoshi Suhara, Viet Anh Trinh, Jinhan Wang, Piotr Zelasko, Hui Wang, Puhui Meng, Chaosen Zhang, Yunsheng Liu, Shawn Wang, Wenjing Li, Zhonglei He