Hugging Face Trending Papers

When Compression Scores Cannot Decide: Information Boundaries for Group-Robust LLM Pruning

A stable compression score can still select the worse model. In our dense study, a split-half reliable path-quadratic score predicted a 16.

arXiv Machine Learning
1d ago

SimplexUQ: An Evaluation Framework and Benchmark for Conformal Uncertainty on Simplex-Valued Predictions

SimplexUQ introduces the first benchmark and reproducible protocol for evaluating how conformal prediction wrappers allocate coverage across simplex‑valued predictions. The framework, called SimplexTasks‑12, combines six synthetic regimes and six real tasks (e.g., class probabilities, topic mixtures, spectral abundances) to compare existing wrappers on metrics such as marginal coverage, worst‑stratum coverage, max disparity, and computational cost. Empirical results show that no single wrapper consistently dominates, with Mondrian and BatchMVP performing best in different settings, and that removing predictor bias only partially mitigates disparity.

By Liang You, Hengyu Shi, Dongwen Ou
arXiv AI
2d ago

False Floors: LLM Safety Routing Evaluations Break Under Distribution Shift

The paper examines safety routers—systems that route user requests to different language models—and finds that their performance degrades significantly when evaluated under distribution shift. In standard benchmarks, routers appear effective because the best single model is chosen from the same evaluation data, but when the data distribution changes, the routing advantage diminishes or disappears. The study quantifies this bias across multiple safety corpora, showing that routers offer little benefit under realistic shift conditions and that recognition‑based defenses can be undermined by attackers who know the model being used.

By Amit Singh Bhatti, Vishal Vaddina
arXiv AI
Jun 29

When Is an LLM Worth It for Hyperparameter Optimization? A Budget-Matched Study on Tabular Data Finds the Warm-Start Is a Default Configuration, Not the Model

arXiv:2606. 21641v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have been proposed as hyperparameter-optimization (HPO) advisors that "warm-start" search from prior knowledge, proposing strong configurations in very few evaluations.

By Carson Rodrigues, Oysturn Vas, Isaiah Abner DCosta, Nithish Kumar Prabhakaran
arXiv Computer Vision
Sep 18

Should This Case Be Adapted? Prediction Fragmentation Controls Test-Time Adaptation

The paper introduces a method for deciding whether to adapt a frozen segmentation model at test time, arguing that a fixed adaptation horizon conflates two distinct decisions: how far to adapt and whether to adapt at all. By measuring disagreement geometry—called prediction fragmentation—between the source model and the adapted mask, the authors predict harmful accepted area (HA) without extra labels or backward passes, achieving strong correlation across three medical benchmarks. A case‑level router built on this metric reduces HA significantly while maintaining or improving Dice scores, and the approach generalizes across architectures and domains.

By Lili Wang, Jing Li, Xiaowen Sun, Xiangyu Hu, Zhuangzhuang Gu, Jian Liu, Srihari Nelakuditi, Yan Tong
arXiv AI
Sep 15

When Should a World Model Move? Loss-Conditioned State Execution

The paper introduces loss‑conditioned state execution, a model‑agnostic technique that decides whether to apply a world model’s proposed state change or keep the current state based on whether the change reduces downstream loss. It formalizes state movability as the existence of a loss‑reducing feasible correction and constructs loss‑specific proposals from predictive distributions, executing them only when a groupwise lower confidence bound on loss improvement is positive. Experiments on forecasting and dynamics benchmarks show that the method accepts updates for a subset of cases, achieving lower bounded loss than persistence or always executing the proposal, and highlights that event predictability and loss‑based decisions must be evaluated separately.

By Jintao Xu, Zhengyu Chen, Ben Zhang, Yongzhi Qi, Jianshen Zhang
Hugging Face Trending Papers
Sep 17

Should This Case Be Adapted? Prediction Fragmentation Controls Test-Time Adaptation

The paper investigates test‑time adaptation for medical image segmentation, showing that a fixed adaptation horizon can harm many individual cases. It introduces prediction fragmentation—a measure of disagreement between the source model and the adapted mask—to predict harmful adaptation without extra labels or backward passes. Using a case‑level router based on this metric, the authors reduce harmful adaptation on cardiac MRI from 58.7% to 20% while maintaining accuracy.