When Compression Scores Cannot Decide: Information Boundaries for Group-Robust LLM Pruning
A stable compression score can still select the worse model. In our dense study, a split-half reliable path-quadratic score predicted a 16.
arXiv:2608. 02940v1 Announce Type: new Abstract: A reproducible compression statistic can still select the wrong candidate.
A stable compression score can still select the worse model. In our dense study, a split-half reliable path-quadratic score predicted a 16.
arXiv:2608.29765v1 Announce Type: new Abstract: Structured pruning uses surrogate objectives because direct task evaluation over every feasible mask is too expensive. Most evaluations report average...
arXiv:2607. 27275v1 Announce Type: new Abstract: Post-training quantization to 4-bit weights is widely reported to be nearly lossless.
arXiv:2608.21382v1 Announce Type: new Abstract: Multiple-choice benchmarks fix the questions and the correct answers, but not the harness: the order of the options, the wording of the prompt, and whe...
arXiv:2608. 06564v1 Announce Type: new Abstract: Quantization is how large language models are actually deployed, and below four bits it is known to hurt.
arXiv:2608. 06564v2 Announce Type: replace Abstract: Quantization is how large language models are actually deployed, and below four bits it hurts.
The paper investigates test‑time adaptation for medical image segmentation, showing that a fixed adaptation horizon can harm many individual cases. It introduces prediction fragmentation—a measure of disagreement between the source model and the adapted mask—to predict harmful adaptation without extra labels or backward passes. Using a case‑level router based on this metric, the authors reduce harmful adaptation on cardiac MRI from 58.7% to 20% while maintaining accuracy.
arXiv:2608. 08652v1 Announce Type: cross Abstract: We present \LegoLM{}, a structured weight-sharing compression framework for large language models grounded in a systematic study of why global weight sharing fails and how to fix it.
arXiv:2608.23744v1 Announce Type: new Abstract: Split conformal prediction, not the pruning rule, supplies finite-sample marginal coverage once a pruned model is fixed independently of the conformal...
The paper introduces a method for deciding whether to adapt a frozen segmentation model at test time, arguing that a fixed adaptation horizon conflates two distinct decisions: how far to adapt and whether to adapt at all. By measuring disagreement geometry—called prediction fragmentation—between the source model and the adapted mask, the authors predict harmful accepted area (HA) without extra labels or backward passes, achieving strong correlation across three medical benchmarks. A case‑level router built on this metric reduces HA significantly while maintaining or improving Dice scores, and the approach generalizes across architectures and domains.
arXiv:2607. 16721v1 Announce Type: new Abstract: The strongest open-weight coding models are mixture-of-experts (MoE) networks: most of their size comes from large pools of "expert" subnetworks, of which only a few act on any token.
The paper introduces Global Relative Kinetic Utility (Global RKU), a label‑free method for calibrating cross‑layer credit in global structured pruning of large language models. Global RKU estimates channel importance via a final‑hidden‑state activation‑gradient signal and applies block‑relative normalization to remove block‑common scale while preserving within‑block ordering, enabling a single‑stage static pruning topology. Experiments on Qwen‑2.5‑7B show significant performance gains at various sparsity levels, and ablation studies confirm the effectiveness of the relative‑normalization step.