arXiv Machine Learning

GeoRVQ: Decoder-aware geometry for residual-token prediction in physiological signals

GeoRVQ introduces a coarse‑to‑fine masked token model that incorporates decoder‑induced geometry into residual‑vector‑quantized physiological waveforms. By using geometry‑aware soft targets and expected distortion, the model improves exact token accuracy from 0.133 to 0.143, reduces decoded distance from 0.606 to 0.393, and raises R‑peak F1 from 0.784 to 0.837 on datasets such as MIMIC‑IV Waveform, VitalDB, and CODE‑15%. The approach demonstrates that decoder‑aware objectives can enhance waveform and event preservation without a large increase in token accuracy.

arXiv AI
6d ago

Audio Token Attention Is Predictable Before the Language Model Runs

The paper introduces Triage, a method that predicts the attention distribution of audio tokens before a language model processes them, enabling early pruning of less important tokens. By fitting a linear map to encoder outputs, Triage achieves high correlation (ρ ≥ 0.69) with full-model attention across eleven of thirteen large audio language models. Using this prediction, Triage compresses audio inputs while maintaining near‑full performance, outperforming baselines in transcription accuracy and significantly increasing the amount of audio that fits within a model’s context window.

By Kyoungjun Park, Yunzhe Li, Lili Qiu
arXiv Machine Learning
Sep 14

TokenMapper: A Step Toward Interoperable Speech Token Translation

TokenMapper is a framework that enables direct translation between different speech tokenizers, allowing heterogeneous speech models to communicate without converting tokens to waveform audio. It handles mismatched token spaces, including single and multi-codebook representations, while maintaining a shared effective token rate. Experiments on GLM-4-Voice, MiMi, and DualCodec show that TokenMapper achieves word error rates close to native reconstructions, comparable human MOS scores, and significantly reduces latency compared to waveform bridging.

By Tal Kozakov, Tal Rosenwein, Eliya Nachmani
arXiv Machine Learning
Sep 24

Six Layers Less: Encoder Pruning for Whisper with Label-Free Recovery

The paper introduces a method to prune six layers from the encoder of OpenAI’s Whisper ASR model, reducing the encoder stack by 18.5% without requiring custom inference code. Layers are selected based on their minimal impact on Word Error Rate when removed. After pruning, the model’s WER rises from 18.2% to 21.9%, but distillation with unlabeled monolingual speech data lowers it to 20.1%. "whyItMatters":"The approach offers a straightforward way to accelerate Whisper inference by simplifying the encoder while maintaining acceptable accuracy, and the released code and model enable immediate adoption by the community."

By Rasmus Aagaard, Nicki Skafte Detlefsen
arXiv Machine Learning
Aug 18

The Null Token Knows: Reducing Message-Free Hallucination in ASR and NMT

The paper investigates how encoder-decoder models in ASR and NMT can generate fluent text even when the input contains no recoverable message, a phenomenon known as message-free hallucination. By auditing the models’ reserved null tokens and manipulating their scores, the authors show that a higher null-token score can suppress fabrication but may also delete valid content or shorten translations. The study highlights that the null token can serve as a diagnostic tool for hallucination and suggests evaluating abstention methods by considering both suppression and deletion costs.

By Kirill Borodin, Vasiliy Kudryavtsev, Ivan Viakhirev
arXiv AI
Sep 28

Softmax Reparameterization for Output-Head Quantization

The paper introduces a post‑training softmax reparameterization technique that selects a functionally equivalent output head before quantization. By subtracting a scalar multiple of the vocabulary‑row mean from each output row and tuning this coefficient via validation KL, the method preserves the full‑precision softmax distribution while enabling efficient W4 quantization. Experiments on seven heads show significant error reductions and latency improvements, with the approach remaining complementary to other quantization strategies and transferable across datasets.

By Asim Kadav, Christian Flores, Chirag Arora, Varun Kotte, Hongbo Zheng, Lan Yan, Priya Shanmugasundaram, Tracy Holloway King