Capability Scaling-Down Laws for LLM Compression
Read the original on arXiv Machine Learning →The paper presents a systematic study of how different compression techniques—pruning, quantization, and distillation—affect the capabilities of large language models (LLMs) in tasks such as mathematics, code generation, and question answering. It introduces a framework that measures capability loss and relates it to factors like model size, training stage, and compression settings, yielding simple predictive relations that generalize across unseen configurations. The authors demonstrate that sharing density responses across pruning levels can dramatically reduce the number of measurements needed, and that their predictive models closely match regression results while offering efficient decision guidance for compression method selection.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.