arXiv Machine Learning By Bo Cui, Yaowen Zhang

GeoRVQ: Decoder-aware geometry for residual-token prediction in physiological signals

Read the original on arXiv Machine Learning →

GeoRVQ introduces a coarse‑to‑fine masked token model that incorporates decoder‑induced geometry into residual‑vector‑quantized physiological waveforms. By using geometry‑aware soft targets and expected distortion, the model improves exact token accuracy from 0.133 to 0.143, reduces decoded distance from 0.606 to 0.393, and raises R‑peak F1 from 0.784 to 0.837 on datasets such as MIMIC‑IV Waveform, VitalDB, and CODE‑15%. The approach demonstrates that decoder‑aware objectives can enhance waveform and event preservation without a large increase in token accuracy.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
6d ago

Audio Token Attention Is Predictable Before the Language Model Runs

The paper introduces Triage, a method that predicts the attention distribution of audio tokens before a language model processes them, enabling early pruning of less important tokens. By fitting a linear map to encoder outputs, Triage achieves high correlation (ρ ≥ 0.69) with full-model attention across eleven of thirteen large audio language models. Using this prediction, Triage compresses audio inputs while maintaining near‑full performance, outperforming baselines in transcription accuracy and significantly increasing the amount of audio that fits within a model’s context window.

By Kyoungjun Park, Yunzhe Li, Lili Qiu
arXiv Machine Learning
Sep 14

TokenMapper: A Step Toward Interoperable Speech Token Translation

TokenMapper is a framework that enables direct translation between different speech tokenizers, allowing heterogeneous speech models to communicate without converting tokens to waveform audio. It handles mismatched token spaces, including single and multi-codebook representations, while maintaining a shared effective token rate. Experiments on GLM-4-Voice, MiMi, and DualCodec show that TokenMapper achieves word error rates close to native reconstructions, comparable human MOS scores, and significantly reduces latency compared to waveform bridging.

By Tal Kozakov, Tal Rosenwein, Eliya Nachmani
arXiv Machine Learning
Sep 24

Six Layers Less: Encoder Pruning for Whisper with Label-Free Recovery

The paper introduces a method to prune six layers from the encoder of OpenAI’s Whisper ASR model, reducing the encoder stack by 18.5% without requiring custom inference code. Layers are selected based on their minimal impact on Word Error Rate when removed. After pruning, the model’s WER rises from 18.2% to 21.9%, but distillation with unlabeled monolingual speech data lowers it to 20.1%. "whyItMatters":"The approach offers a straightforward way to accelerate Whisper inference by simplifying the encoder while maintaining acceptable accuracy, and the released code and model enable immediate adoption by the community."

By Rasmus Aagaard, Nicki Skafte Detlefsen