arXiv:2608. 15046v1 Announce Type: new Abstract: A fraction of a point of benchmark accuracy is the usual evidence that a compressed model is equivalent to its original.
By Amogh Singh
The study analyzes 373,019 judgments from LLM‑scored benchmarks, decomposing variance into system, item, judge, and interaction components via generalizability theory. It finds that with a single judge, generalizability converges to a ceiling determined by the system‑by‑judge variance, which is substantially lower in pairwise preference settings, allowing one judge to suffice. The research also reveals significant biases in presentation order and highlights that many published win‑rate claims fall below the measured floor of the benchmarks.
By Atul Anand
The paper introduces a calibrated instrument for rigorously measuring how inference optimizations—such as quantization, early‑exit, and speculative decoding—affect the output quality of large language models. It uses a formally calibrated LLM judge that verifies no systematic bias between statistically equivalent outputs and includes a null condition to ensure measured differences are zero. Applying this method, the authors find that a 4‑bit model is indistinguishable from its 16‑bit counterpart, while 3‑bit quantization and early‑exit techniques incur measurable quality losses that vary by language and task, and that token‑certainty‑based acceptance rules cannot reliably identify impactful errors.
By Jerry Kaplan
arXiv:2609.07944v1 Announce Type: new
Abstract: Existing causal-inference benchmarks for LLMs mostly score method descriptions or whether generated code runs, not whether the executed workflow recove...
By Yonghong Zhang, Ricardo Correia, Isabel M. Parra, Yong Xie
Large language model optimization is an active research area, spanning quantization of model weights, early-exit methods for skipping layers, and speculative decoding. Each track uses its own quality...
The paper investigates how sequential knowledge editing can degrade a language model’s ability to discern reliable evidence from unreliable evidence without affecting overall accuracy. Using a conservatively tuned LoRA on Qwen2.5‑7B‑Instruct, the authors show that after 1,000 edits the model’s arbitration score for untouched facts drops by 36%, leading to higher error rates on its most confident decisions, while MMLU accuracy remains unchanged. The study also finds that in some model‑method combinations, sequential edits can reduce MMLU to chance levels even though edit success and locality remain perfect.
By Atul Anand