arXiv:2607. 16821v1 Announce Type: cross Abstract: Task arithmetic, sequential fine-tuning, activation steering, and first-order random search all operate through relatively small perturbations around an already trained checkpoint, and they rely on different local approximations: individual perturbations should be first-order predictable, task updates should compose with controlled interference, useful tangent structure should be stable and possible to estimate, and weight edits should have counterparts in representation space.
By Irina Piontkovskaia, Sergey Nikolenko
arXiv:2608. 06206v1 Announce Type: cross Abstract: Conformal prediction endows arbitrary black-box predictors with finite-sample, distribution-free marginal coverage, yet marginal validity can hide severe covariate-specific miscalibration, while exact distribution-free conditional coverage is finite-sample unattainable.
By Anton Conrad, Rustam Isaev, Denis Belomestny, Eric Moulines, Sergey Samsonov
arXiv:2609.37858v1 Announce Type: new
Abstract: Many localized large language model (LLM) unlearning methods select a small parameter subset from a localization signal and keep it fixed during optimi...
By Tianhao Qian, Ziming Hong, Chongyang Gao, Kezhen Chen, Lixu Wang
arXiv:2607. 27680v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA) has become the standard mechanism for fine-tuning large pretrained models, yet its statistical properties remain only partially understood.
By Arunan J
arXiv:2608.23744v1 Announce Type: new
Abstract: Split conformal prediction, not the pruning rule, supplies finite-sample marginal coverage once a pruned model is fixed independently of the conformal...
By Ibne Farabi Shihab, Adria Binte Habib, Anuj Sharma
arXiv:2608. 01460v1 Announce Type: new Abstract: Conformal prediction (CP) is a distribution-free framework for uncertainty quantification that has recently been adapted to large language models (LLMs), providing prediction sets with finite-sample coverage guarantees under exchangeability.
By Yuqicheng Zhu, Jialin Yu, Lin Li, Gengyuan Zhang, Zhen Yang, Steffen Staab, Puneet Dokania, Philip Torr, Jie Tang, Evgeny Kharlamov
arXiv:2607. 18088v1 Announce Type: new Abstract: Standard evaluation of many recognition systems contains distribution shift by construction, since benchmarks place disjoint conditions in the training and test splits.
By Weijia Han, Lisha Qu
arXiv:2606. 00257v1 Announce Type: cross Abstract: Token-level credit assignment for language-model reinforcement learning is usually formulated as if the policy were fully trainable, while practical LLM-RL pipelines often rely on parameter-efficient fine-tuning, especially LoRA.
By Rodney Lafuente-Mercado
arXiv:2609.38917v1 Announce Type: new
Abstract: A classifier's conditional accuracy can change while its confidence distribution stays exactly the same. We study the worst-case movement of the reliab...
By Wenhao Liang, Lin Yue, Wei Emma Zhang, Mingyu Guo, Olaf Maennel, Weitong Chen
arXiv:2609.10333v1 Announce Type: new
Abstract: Uncertainty estimation for medical vision--language models (VLMs) using conformal prediction has gained increasing attention due to its distribution-fr...
By Xuan Cuong Ngo, Ngan Le
DeepGOF-1 introduces a pretrained convolutional network as a goodness‑of‑fit test for logistic regression, where the network reads a grid of standardized residuals as an image and outputs a test statistic. The test is fully calibrated via the analyst’s own bootstrap, ensuring the nominal level is maintained regardless of the network’s training. The authors prove exactness under pivotality, asymptotic exactness without it, and provide a computable consistency certificate from the frozen weights, demonstrating superior stability and power across multiple benchmarks and sample sizes.
By Ebrahim Khaled Ebrahim
arXiv:2607. 12501v3 Announce Type: replace Abstract: The Forward-Forward algorithm trains each layer locally, so that a scalar goodness - the sum of squared activations - is high on real inputs and low on contrastive ones.
By Paolo Giannitrapani