arXiv AI

The Shape of Addition: Geometric Structures of Arithmetic in Large Language Models

arXiv:2606. 03645v1 Announce Type: cross Abstract: Large Language Models exhibit paradoxical fragility in fundamental arithmetic, implying a disconnect between internal computation and discrete output.

arXiv Machine Learning
Sep 11

Why Does Post-Training Quantization Work?

Post‑training quantization compresses large language models by storing weights at reduced precision, introducing errors into hidden states that could accumulate with depth. However, pretrained models accumulate far less hidden‑state error than randomly initialized ones, largely preserving downstream performance. The study identifies two key mechanisms: (1) each layer’s new error tends to oppose inherited error, partially canceling it, and (2) the LM‑head geometry preserves high‑rank token scores, mitigating output changes.

By Yuxiang Chen, Michael Beyer, Jun Zhu, Jianfei Chen
Hugging Face Trending Papers
Jun 3

STaR-Quant: State-Time Consistent Post-Training Quantization for Diffusion Large Language Models

Diffusion large language models (DLLMs) have recently emerged as a promising alternative to autoregressive LLMs by generating text through iterative masked denoising with bidirectional context. However, their large model sizes and iterative denoising process introduce substantial memory and computational overhead, motivating post-training quantization for efficient deployment.

arXiv AI
Sep 21

Understanding In-context Learning of Addition via Activation Subspaces

The paper investigates how transformer language models perform few‑shot learning for a simple addition task, showing that the ability is concentrated in a handful of attention heads. Using dimensionality reduction, the authors identify low‑dimensional subspaces—three heads with six‑dimensional spaces in Llama‑3‑8B‑Instruct—where specific dimensions encode the units digit via trigonometric patterns and magnitude via low‑frequency components. They also derive a mathematical identity linking aggregator and extractor subspaces, enabling tracking of information flow from examples to the final prediction.

By Xinyan Hu, Kayo Yin, Michael I. Jordan, Jacob Steinhardt, Lijie Chen