arXiv Machine Learning

Low-Bit Recurrent States in Hybrid Language Models

The paper introduces a method for quantizing the fixed‑size recurrent states of hybrid language models to as few as four bits per token. By deriving distortion weights from the observability Gramian and combining them with normalized state ranges, the authors achieve mixed‑precision bit allocation without requiring calibration data, rotation, or additional training. The approach also logarithmically quantizes decay rates, yielding significant reductions in excess negative log‑likelihood—up to 27.9× better than seven baselines—while maintaining near‑FP32 performance at six bits.

arXiv Machine Learning
Aug 31

DAMP: Decay-Aware Mixed-Precision Recurrent-State Quantization

The paper introduces DAMP, a decay‑aware mixed‑precision quantization scheme for recurrent‑state representations in GDN and KDA language models. By identifying high‑risk channels through quantization‑error energy and decay persistence, DAMP stores these channels at higher precision while compressing the rest to INT8, achieving a 9.9‑bit average precision. Experiments on Qwen3.6‑35B and Kimi‑Linear‑48B show a 69.1% reduction in recurrent‑state storage, up to 2.01× faster state‑update kernels, and up to 10.9% lower full‑model TPOT while preserving accuracy close to the FP32 baseline.

By Tao Zhang, Jianchao Tan, Pingwei Sun, Yanqi Yu, Zixu Jiang, Yuchen Xie, Xunliang Cai, Ziqian Zeng
arXiv Machine Learning
6d ago

Same Bit Width, Different Outcomes: Post-Training Quantization of Text-to-Speech Across Architectures

The paper evaluates post‑training quantization (PTQ) for text‑to‑speech (TTS) models across multiple architectures using a unified protocol. It shows that reducing weights to 4‑bit per‑channel can significantly lower predicted mean opinion scores (UTMOS) and that even 8‑bit per‑tensor scaling can cause severe degradation, with the impact varying by model. A staged ablation identifies the sensitive components, and per‑layer GPTQ can recover performance to within 0.1 UTMOS, while real int8 and int4 kernels confirm the simulated results on hardware, demonstrating that each configuration must be validated on the target runtime.

By Se Un Park, Yutae Kim, Junyoung Park
arXiv Machine Learning
Sep 10

Squeeze10-LLM: Squeezing LLMs' Weights by 10 Times via a Staged Mixed-Precision Quantization Method

Squeeze10-LLM is a staged mixed‑precision post‑training quantization framework that reduces 16‑bit LLM weights to an average of 1.6 bits per weight by assigning 80% of weights to 1 bit and 20% to 4 bits. It introduces Post‑Binarization Activation Robustness (PBAR), a weight significance metric that considers activation impact, and Full Information Activation Supervision (FIAS), a strategy that preserves activation information to limit error propagation. Experiments on LLaMA and LLaMA2 demonstrate that Squeeze10‑LLM achieves state‑of‑the‑art performance for sub‑2‑bit weight‑only quantization, raising average accuracy from 43% to 56% on six zero‑shot classification tasks.

By Qingcheng Zhu, Yangyang Ren, Linlin Yang, Yanjing Li, Sheng Xu, Haodong Zhu, Juan Zhang, Runqi Wang, Baochang Zhang
arXiv Computation and Language
Sep 23

LatentPort: Beyond KV Cache - Cross-Model Transfer of Recurrent Memory in Hybrid Language Models: A 4B-to-9B Hybrid-State Handoff Without Target Prefix Replay

The paper introduces LatentPort, a method that allows a language model to transfer its live memory to another model without requiring the receiver to reread the context. Experiments on a Qwen3.5 4B-to-9B sibling pair show that adding a Gated DeltaNet (GDN) persistent-state package reduces negative log‑likelihood by 0.747 nats/token and improves performance across 64 PG19 documents. The study also demonstrates that direct recurrent and convolution reuse outperforms learned GDN maps, and a 434,176‑parameter correction further narrows the performance gap to the native 9B model.

By Simon P. Villani
arXiv AI
Sep 17

Objective vs. Search: Decomposing What Makes a Good Tokeniser

The paper introduces two new tokenisation algorithms—BottomUpLL and TopDownComp—to systematically explore the 2x2 design space defined by optimisation objective (compression vs. log‑likelihood) and search procedure (bottom‑up merging vs. top‑down pruning). Experiments across model sizes, vocabularies, and domains show that the search procedure, rather than the objective, consistently yields lower bits‑per‑byte, while no clear pattern emerges on the BLiMP benchmark. These findings clarify how tokeniser design choices influence language‑model performance and provide guidance for constructing tokenisers more principledly.

By Ahmetcan Yavuz, Clara Meister, Tiago Pimentel
arXiv Machine Learning
Sep 14

Breaking the Token Ceiling: Distilling Smaller, Stronger Byte Models

The paper investigates whether small models distilled from larger ones behave similarly when using byte versus token tokenization. It introduces two methods—Marginalize‑It (approximate) and End‑Of‑Token (exact)—to convert token logits to byte logits, and conducts a large‑scale study on decoder‑only dense transformers ranging from 1 billion to 1 trillion bytes of data. Results show that while token‑based models excel early, byte‑based models eventually surpass them with more compute, achieving higher performance ceilings, greater data efficiency, and lower logit storage costs.

By Kalyani Marathe, Artidoro Pagnoni, Tomasz Limisiewicz, Margaret Li, Mike Lewis, Luke Zettlemoyer, Srinivasan Iyer
arXiv AI
6d ago

An Exploratory Ablation of a Small MLA--SSM Hybrid Language Model

The paper reports a single‑seed ablation study of the TALH language model, which combines a Multi‑head Latent Attention (MLA) branch with a custom recurrent state‑space model (SSM). Five variants ranging from 117 M to 217 M active parameters per token were trained on a FineWeb sample, and the results show that removing the SSM branch causes the largest drop in validation perplexity (315) compared to removing MLA (239). A dense‑FFN hybrid achieved a perplexity of 231, outperforming the tested top‑2 ternary‑MoE hybrid (240) while using 3.87 GB less peak training memory, and MLA‑only exhibited the flattest time‑to‑first‑token curve on an Apple M3, though the dense Transformer was faster overall.

By Christos Koutsiaris