arXiv Machine Learning By Wietse Stienstra

One Ruler: A Same-Hands Re-Evaluation of Bivariate Causal Direction on Tuebingen, with a Parameter-Free Compression Baseline

Read the original on arXiv Machine Learning →

arXiv:2606. 23767v1 Announce Type: new Abstract: Headline accuracies on the Tuebingen cause-effect pairs are routinely compared across papers even though each is measured under its authors' own protocol -- different pair subsets, weightings, model-selection, and decision rates.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 24

Ask Which, Not How Good: Sizing Benchmarks Scored by an LLM

The study analyzes 373,019 judgments from LLM‑scored benchmarks, decomposing variance into system, item, judge, and interaction components via generalizability theory. It finds that with a single judge, generalizability converges to a ceiling determined by the system‑by‑judge variance, which is substantially lower in pairwise preference settings, allowing one judge to suffice. The research also reveals significant biases in presentation order and highlights that many published win‑rate claims fall below the measured floor of the benchmarks.

By Atul Anand
arXiv Machine Learning
Sep 17

A Calibrated Instrument for Measuring How Inference Optimizations Affect Output Quality

The paper introduces a calibrated instrument for rigorously measuring how inference optimizations—such as quantization, early‑exit, and speculative decoding—affect the output quality of large language models. It uses a formally calibrated LLM judge that verifies no systematic bias between statistically equivalent outputs and includes a null condition to ensure measured differences are zero. Applying this method, the authors find that a 4‑bit model is indistinguishable from its 16‑bit counterpart, while 3‑bit quantization and early‑exit techniques incur measurable quality losses that vary by language and task, and that token‑certainty‑based acceptance rules cannot reliably identify impactful errors.

By Jerry Kaplan
arXiv AI
Sep 25

Sequential knowledge editing breaks a model's ability to tell good evidence from bad, without costing it accuracy

The paper investigates how sequential knowledge editing can degrade a language model’s ability to discern reliable evidence from unreliable evidence without affecting overall accuracy. Using a conservatively tuned LoRA on Qwen2.5‑7B‑Instruct, the authors show that after 1,000 edits the model’s arbitration score for untouched facts drops by 36%, leading to higher error rates on its most confident decisions, while MMLU accuracy remains unchanged. The study also finds that in some model‑method combinations, sequential edits can reduce MMLU to chance levels even though edit success and locality remain perfect.

By Atul Anand