arXiv AI

Offline AI Modules: Voice-First Offline Architecture, Hardware Reference Stack, Quantization and Benchmarking

The paper introduces the Offline AI Modules workstream, which provides a voice‑first offline architecture, a low‑cost hardware reference bill of materials, and a reproducible quantization and benchmarking pipeline for instruction‑tuned language models in the 2‑5B parameter range. It evaluates three models across four quantization formats on two hardware tiers—NVIDIA Jetson Orin NX and Raspberry Pi5—measuring deployment metrics and multilingual quality. The key result is that Q4_K_M quantization offers the best size‑to‑quality trade‑off, enabling high decode throughput and strong topic classification accuracy while staying within memory limits on both tiers.

arXiv Machine Learning
Sep 29

LUMO (Lightweight Unified Multilingual Orchestrator): A Privacy Preserving Offline Voice Assistant

LUMO (Lightweight Unified Multilingual Orchestrator) is a privacy‑preserving offline voice assistant that runs entirely on edge hardware, specifically a Raspberry Pi 5 with 8 GB RAM. It integrates local ASR, a 4‑bit GGUF‑quantized LLM, and TTS to deliver end‑to‑end response latencies of 2.0–4.0 s, a 6.8 % WER on short English utterances, and lower peak power consumption (~9 W) compared to existing edge assistants. The system also supports Bangla speech, enabling multilingual use in low‑resource settings.

By Md. Mehedi Hasan Naeem, Mst. Kamrunnahar Ruma, Nafiza Anjum, Shakila Sultana, Md. Sujan Ali
arXiv Machine Learning
Sep 25

Same Bit Width, Different Outcomes: Post-Training Quantization of Text-to-Speech Across Architectures

The paper evaluates post‑training quantization (PTQ) for text‑to‑speech (TTS) models across multiple architectures using a unified protocol. It shows that reducing weights to 4‑bit per‑channel can significantly lower predicted mean opinion scores (UTMOS) and that even 8‑bit per‑tensor scaling can cause severe degradation, with the impact varying by model. A staged ablation identifies the sensitive components, and per‑layer GPTQ can recover performance to within 0.1 UTMOS, while real int8 and int4 kernels confirm the simulated results on hardware, demonstrating that each configuration must be validated on the target runtime.

By Se Un Park, Yutae Kim, Junyoung Park
arXiv Machine Learning
Sep 10

TontaubeV1: Streaming Text-to-Speech with Hierarchical Codec Modeling and Bounded Context

TontaubeV1 is a streaming text‑to‑speech model that preserves natural prosody while running on a single consumer GPU. It encodes speech with a hierarchical DualCodec representation at 12.5 Hz, separating a semantic stream from successive acoustic refinements, and uses Qwen3‑derived transformers to predict the semantic stream, utterance duration, and acoustic refinements. The system supports up to one minute of reference audio for voice conditioning, streams with a 200 ms latency to first audio, and achieves real‑time factors of 0.08 (single input) and 0.02 (eight concurrent inputs).

By Fritz Cremer, Jonathan Cremer
Hugging Face Trending Papers
Sep 8

TontaubeV1: Streaming Text-to-Speech with Hierarchical Codec Modeling and Bounded Context

TontaubeV1 is a streaming text‑to‑speech model that maintains natural prosody while running on a single consumer GPU. It encodes speech with a hierarchical DualCodec representation at 12.5 Hz, separating a semantic stream from successive acoustic refinements, and uses Qwen3‑derived transformers to predict the semantic stream and add refinements. The model supports up to one minute of reference audio for voice conditioning, streams with a 200 ms latency to first audio, and achieves real‑time factors of 0.08 (single input) and 0.02 (eight concurrent inputs).

arXiv AI
Sep 21

NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities

arXiv:2609.21967v1 Announce Type: cross Abstract: We introduce NemotronLabs VoiceChat, an open full-duplex speech-to-speech model with native tool-calling capabilities. NemotronLabs VoiceChat combine...

By Jagadeesh Balam, Travis Bartley, Edresson Casanova, Sanjay Chauhan, Chen Chen, Zhehuai Chen, Zijia Chen, Francesco Ciannella, Slyne Deng, Mikyas Desta, Harishchandra Dubey, Slim Essid, Nourchene Ferchichi, Boris Ginsburg, Mariana Graterol Fuenmayor, Negar Habibi, Kevin Hu, Anand Joseph, Viraj Karandikar, Myungjong Kim, Viacheslav Klimkov, Seelan Lakshmi Narasimhan, Lily Lee, Jason Li, Eileen Long, Ameya Mahabaleshwarkar, Aditya Malte, Adi Margolin, Sasha Meister, Valentin Mendelev, Oluwatobi Olabiyi, Ankita Pasad, Yifan Peng, Elena Rastorgueva, Jayda Ritchie, Jason Roche, Nikhil Srihari, Yuanhang Su, Yoshi Suhara, Viet Anh Trinh, Jinhan Wang, Piotr Zelasko, Hui Wang, Puhui Meng, Chaosen Zhang, Yunsheng Liu, Shawn Wang, Wenjing Li, Zhonglei He
arXiv Computation and Language
Sep 24

Text Scores Can Miss Waveform Use: A Qwen2-Audio Quantization Case Study

The paper presents a new evaluation protocol for post‑training quantization of speech language models that separates lexical output, transcript‑insufficient endpoints, and packed implementations. In a Qwen2‑Audio case study, a 6‑bit allocation selected for translation improves chrF scores but degrades emotion recognition, while uniform and front‑layer controls perform better on emotion tasks. Similar patterns hold at 7 bits, and a 4‑bit study shows consistent emotion deficits across all low‑bit allocations, with no advantage for the selected scheme. The study highlights a precision‑dependent mismatch between lexical output, waveform‑dependent behavior, and nominal precision, without claiming a general failure of low‑bit models or a deployment benefit for the selected allocation.

By Mengzhe Geng, Jinxi Jin, Junhao Xu
arXiv AI
Jul 29

How Small Can You Go? A Controlled Study of LoRA Rank, Target Modules, and Quantization Trade-offs for Text-to-SQL on a 60M-Parameter Model

arXiv:2607. 25583v1 Announce Type: new Abstract: Parameter-efficient fine-tuning (PEFT) and low-bit quantization are now standard tools for adapting language models under tight compute budgets, yet their interaction is most often studied on billion-parameter models where the design space is expensive to explore.

By Mahendra Singh Rathor, Anagheem Azzam
arXiv Machine Learning
Sep 17

Beyond Static RAG: An Adaptive, Tri-Metric Routing Framework for Efficient Long-Context Inference on Commodity GPUs

The paper introduces the Tri‑Metric Router, a deterministic, training‑free policy that chooses among Raw, Neural, and Lexical pipelines for retrieval‑augmented generation on commodity GPUs. It uses three CPU‑side signals—spatial complexity, syntactic density, and type‑token ratio—to balance VRAM headroom and latency, calibrated on LongBench qasper. The method eliminates out‑of‑memory failures and improves alignment and F1 scores compared to always‑on lexical compression without extra VRAM or training costs.

By Saipraveen Vabbilisetty, Ajay Kumar Boddepalli, Deep Narayan Mishra, Shashank Kapadia, Haoan Wang, Anupriya Sharma
arXiv AI
6d ago

Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing

The paper presents a billing‑aware neural text‑to‑speech system for serverless CPUs, focusing on minimizing CPU‑seconds and GB‑seconds rather than just throughput or latency. By using request‑sized concurrent inference and a reclaimable instance lifecycle, the system limits per‑request CPU parallelism and releases idle memory while keeping the server process alive. On the Kokoro‑82M benchmark, it achieves 2.71 audio‑seconds per CPU‑second versus 0.89 with ONNX Runtime, cuts cost per audio‑hour from $0.0631 to $0.0153, and reduces idle billed memory from 8.7 GB to 1.33 GB, with faster restoration times.

By Pakorn Nathong, Kunat Pipatanakul