arXiv Computation and Language
Sep 2

Toppling the Hierarchy in Byte-level Language Modeling

The paper investigates why current byte‑level language models, which use a hierarchical structure that down‑samples to words and then upsamples back to bytes, struggle with precise character manipulation. Experiments show that pure byte‑level models outperform hierarchical variants on character‑level tasks, and that byte‑level attention is the key component driving this advantage. The study explains the trade‑off between computational efficiency and fine‑grained character understanding in hierarchical byte models.

By Lukas Edman, Alexander Fraser
arXiv AI
Aug 25

Lexical Perturbations Disrupt LLM Reasoning: An Empirical Study of Attention Diversion

The study investigates how lexical perturbations—such as keyboard noise, character swaps, and filler insertion—affect large language models (LLMs) on reasoning benchmarks. Four open-weight instruction-tuned models and frontier models were evaluated, revealing that character-level perturbations significantly reduce accuracy, especially on multi-step reasoning tasks, while filler insertion has minimal impact. The authors attribute this asymmetry to Attention Diversion, where fragmented subword tokenization draws disproportionate attention in middle and final transformer layers; they demonstrate that both token content and attention allocation are coupled, making it difficult for inference-time repair strategies to fully recover performance.

By Jiaqian Zhu, Yang Zhang, Junhua Ding, Xiaowei Yu