arXiv AI By Kyoungjun Park, Yunzhe Li, Lili Qiu

Audio Token Attention Is Predictable Before the Language Model Runs

Read the original on arXiv AI →

The paper introduces Triage, a method that predicts the attention distribution of audio tokens before a language model processes them, enabling early pruning of less important tokens. By fitting a linear map to encoder outputs, Triage achieves high correlation (ρ ≥ 0.69) with full-model attention across eleven of thirteen large audio language models. Using this prediction, Triage compresses audio inputs while maintaining near‑full performance, outperforming baselines in transcription accuracy and significantly increasing the amount of audio that fits within a model’s context window.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 7

OmniFocus: Query-Guided Modality-Balanced Token Compression for Omni-Modal Large Language Models

arXiv:2607. 03050v1 Announce Type: cross Abstract: Omni modal large language models (OmniLLMs) have attracted wide attention for their ability to jointly process audio and video, but they generate large token sequences under audio-visual inputs, leading to substantial inference cost.

By Shijie Cao, Qingyu Zhang, Boxi Yu, Yuzhong Zhang, Boxi Cao, Yaojie Lu, Hongyu Lin, Xianpei Han, Le Sun
arXiv Machine Learning
Sep 10

TontaubeV1: Streaming Text-to-Speech with Hierarchical Codec Modeling and Bounded Context

TontaubeV1 is a streaming text‑to‑speech model that preserves natural prosody while running on a single consumer GPU. It encodes speech with a hierarchical DualCodec representation at 12.5 Hz, separating a semantic stream from successive acoustic refinements, and uses Qwen3‑derived transformers to predict the semantic stream, utterance duration, and acoustic refinements. The system supports up to one minute of reference audio for voice conditioning, streams with a 200 ms latency to first audio, and achieves real‑time factors of 0.08 (single input) and 0.02 (eight concurrent inputs).

By Fritz Cremer, Jonathan Cremer
arXiv AI
Aug 25

Pre-Decoding Acoustic Triage for Budgeted Vision-Language Captioning of Untrimmed Egocentric Video

The paper introduces an audio-first triage method for budgeted vision‑language captioning of untrimmed egocentric video. By selecting windows for a vision‑language model based on lightweight audio features before any video decoding, the approach reduces costly model calls. It achieves 4.0–10.8 percentage point improvements in action coverage across call rates, cuts 9–20% of VLM calls at matched coverage on EPIC‑KITCHENS‑100, and outperforms uniform sampling and recent visual keyframe selectors on Ego4D.

By Masoud Jalayer, Changyi Li, Yu Xiao
Hugging Face Trending Papers
Sep 8

TontaubeV1: Streaming Text-to-Speech with Hierarchical Codec Modeling and Bounded Context

TontaubeV1 is a streaming text‑to‑speech model that maintains natural prosody while running on a single consumer GPU. It encodes speech with a hierarchical DualCodec representation at 12.5 Hz, separating a semantic stream from successive acoustic refinements, and uses Qwen3‑derived transformers to predict the semantic stream and add refinements. The model supports up to one minute of reference audio for voice conditioning, streams with a 200 ms latency to first audio, and achieves real‑time factors of 0.08 (single input) and 0.02 (eight concurrent inputs).