arXiv Computer Vision

OmniKVQuant: KV Cache Quantization for Omni-LLMs

OmniKVQuant introduces a training‑free framework for quantizing the key‑value (KV) cache of omni‑modal large language models (Omni‑LLMs) that process audio, video, and text simultaneously. The method addresses two identified problems—temporal key drift and heterogeneous value geometry—by setting key quantization ranges over short input windows and rotating values separately for each modality. Applied to Qwen2.5‑Omni and Qwen3‑Omni, OmniKVQuant achieves 2‑bit KV caches while largely preserving performance across seven audio‑visual benchmarks, and includes a fused Triton decode kernel that eliminates the need for a dense FP16 cache.

arXiv AI
1d ago

HyQuant: Hybrid-Precision Quantization for LLM Attention

HyQuant introduces a hybrid-precision quantization framework for large language model (LLM) attention modules. It quantizes most attention states to low-bit formats while preserving a small set of vertical‑line tokens and local‑window states in full precision, guided by lightweight attention‑pattern signals. This design achieves near‑lossless accuracy across tasks while improving memory and hardware efficiency.

By Jiatong Ding, Bingxin Xing, Yu Zhang, Dian Ding, Xiaodong Yi, Xianbin Ouyang, Feihu Zhou, Kun Zhang, Zhenyu Guo, Hao Pan, Guangtao Xue, Yiming Zhang
arXiv AI
Jul 7

OmniFocus: Query-Guided Modality-Balanced Token Compression for Omni-Modal Large Language Models

arXiv:2607. 03050v1 Announce Type: cross Abstract: Omni modal large language models (OmniLLMs) have attracted wide attention for their ability to jointly process audio and video, but they generate large token sequences under audio-visual inputs, leading to substantial inference cost.

By Shijie Cao, Qingyu Zhang, Boxi Yu, Yuzhong Zhang, Boxi Cao, Yaojie Lu, Hongyu Lin, Xianpei Han, Le Sun
arXiv AI
Jul 29

OmniDelta: Skill-Driven Budget Allocation for Token Compression in OmniLLMs

arXiv:2607. 25669v1 Announce Type: new Abstract: Emerging Omni-modal Large Language Models (OmniLLMs) enable unified understanding of text, audio, and video, but their long audio-video token sequences introduce substantial memory and inference costs.

By Haoyang Huang, Wenjie Huang, Tianqi Xu, Hongyaoxing Gu, Kang Tan, Yikai Fu, Yuhao Shen, Tianyu Liu, Baolin Zhang, Jun Zhang, Xinyi Hu, Jun Dai, Shuang Ge, Lei Chen, Yue Li, Mingchen Wang, Meng Zhang