arXiv:2609.01072v1 Announce Type: new
Abstract: Post-hoc calibration corrects reported confidence, yet a multiclass calibrator can also change the associated top-1 prediction. Accuracy captures only...
By Daehwan Kim, Haejun Chung, Ikbeom Jang
arXiv:2609.36721v1 Announce Type: new
Abstract: Fixed-budget adaptation from heterogeneous data sources requires deciding not only how much data to use, but how much exposure each source receives. Si...
By Yupeng Chang, Yuan Wu
The paper introduces RouteCert, a method for ensuring risk control in multimodal systems that acquire inputs adaptively. It shows that conditional calibration can remain valid even when the acquisition policy determines the calibration group, and provides two finite‑sample constructions: threshold‑free routing with terminal‑pattern calibration and simultaneous validation of policy‑pattern pairs. Experiments on a clinical ECG task and masked multimodal benchmarks demonstrate that RouteCert achieves low disagreement rates and competitive answered fractions while validating each acquisition stage separately.
By Melika Baghi
arXiv:2607. 07745v1 Announce Type: new Abstract: While accuracy, robustness, and calibration are all essential for reliable neural networks, they are often studied separately; developing models that satisfy all three simultaneously remains a central challenge.
By Arthur Chiron (IRIT, EPE UT), Franck Mamalet (IRIT, DTIPG - SNCF, UT3), Thomas Massena (IRIT, DTIPG - SNCF, UT3), Thomas Deltort (IRIT), Mathieu Serrurier (IRIT, UT2J)
arXiv:2607. 18162v1 Announce Type: new Abstract: The soft-label Bayes-error estimator beta(z) = E[min(z, 1-z)] of Ishida et al.
By Shreyas Pradeepkumar Khandale
arXiv:2608. 15520v1 Announce Type: new Abstract: A multimodal system may begin inference holding only some of its inputs and may acquire the rest at a cost.
By Melika Baghi
arXiv:2609.38917v1 Announce Type: new
Abstract: A classifier's conditional accuracy can change while its confidence distribution stays exactly the same. We study the worst-case movement of the reliab...
By Wenhao Liang, Lin Yue, Wei Emma Zhang, Mingyu Guo, Olaf Maennel, Weitong Chen
arXiv:2606. 20364v1 Announce Type: new Abstract: A companion study established a de-biased, cross-model VLM-as-3D-judge that reliably ranks single-image-to-3D mesh quality where cheap geometry and CLIP proxies fall short.
By Ali Asaria, Tony Salomone, Deep Gandhi
arXiv:2609.17386v1 Announce Type: new
Abstract: Test-time prompt tuning (TPT) enables adaptation on a single test instance, achieving improved accuracy but often sacrificing calibration performance....
By Yuwei Liang, Jian Liang, Dapeng Hu, Yinuo Xu, Ran He
The paper investigates converting a large pretrained transformer (1.4 B parameters) into a smaller sibling (410 M) by studying representation alignment and parameter projection. It finds that dense weight projection destroys structure, and that a low‑budget, structure‑aware compensation—separating least‑squares function alignment from variance‑preserving rescaling—yields significant gains on token‑efficient training, outperforming subcloning and standard distillation pipelines at matched budgets.
By Ravi Satya Durga Prasad Yenugula
The paper investigates test‑time adaptation for medical image segmentation, showing that a fixed adaptation horizon can harm many individual cases. It introduces prediction fragmentation—a measure of disagreement between the source model and the adapted mask—to predict harmful adaptation without extra labels or backward passes. Using a case‑level router based on this metric, the authors reduce harmful adaptation on cardiac MRI from 58.7% to 20% while maintaining accuracy.
The paper introduces a method for deciding whether to adapt a frozen segmentation model at test time, arguing that a fixed adaptation horizon conflates two distinct decisions: how far to adapt and whether to adapt at all. By measuring disagreement geometry—called prediction fragmentation—between the source model and the adapted mask, the authors predict harmful accepted area (HA) without extra labels or backward passes, achieving strong correlation across three medical benchmarks. A case‑level router built on this metric reduces HA significantly while maintaining or improving Dice scores, and the approach generalizes across architectures and domains.
By Lili Wang, Jing Li, Xiaowen Sun, Xiangyu Hu, Zhuangzhuang Gu, Jian Liu, Srihari Nelakuditi, Yan Tong