arXiv:2607. 12501v3 Announce Type: replace Abstract: The Forward-Forward algorithm trains each layer locally, so that a scalar goodness - the sum of squared activations - is high on real inputs and low on contrastive ones.
By Paolo Giannitrapani
The study investigates whether the number of discrete class‑separability jumps (phase transitions) observed during ResNet fine‑tuning can predict final test accuracy. Across 75 experiments on four benchmarks (CIFAR‑10, CIFAR‑100, TinyImageNet, CIFAR‑10‑C) and three ResNet variants, a strong negative correlation is found on standard i.i.d. datasets (r = −0.84 on CIFAR‑10, r = −0.87 on CIFAR‑100), while the correlation weakens under distributional stress. Additional analyses show that the transition count retains predictive power after controlling for architecture depth and outperforms other training‑curve signals on in‑distribution benchmarks, though it is dominated by other signals on stressed datasets.
By Arunan J
arXiv:2608. 08424v1 Announce Type: cross Abstract: Conformal changepoint localization turns any score into a confidence set for the changepoint with finite-sample coverage.
By Chenchen Peng, Mixia Wu, Qijing Yan, Zhiqi Shen, Jie Zhang
arXiv:2609.39229v1 Announce Type: cross
Abstract: Automatic evaluation of faithfulness increasingly relies on a large language model acting as a judge, yet the most reliable judges are proprietary fr...
By Elia Onofri, Roberto Di Pietro
arXiv:2606. 03305v1 Announce Type: new Abstract: Benchmark contamination, where evaluation examples appear in a model's training data, threatens the validity of LLM assessment.
By Wojciech Zarzecki, Jan Dubi\'nski, Sebastian Cygert
arXiv:2608. 09768v1 Announce Type: new Abstract: A prediction that is both confident and wrong is a critical reliability failure because it can bypass abstention and human review precisely when the model is mistaken.
By Ange-Cl\'ement Akazan, Ineza Remy Mugenga, Abebe Geletu, Jean Medard Ngnotchouye, Issa Karambal
arXiv:2608. 12652v1 Announce Type: cross Abstract: Benchmark contamination is diagnosed today with n-gram overlap, with likelihood-based membership inference, or with canary strings, and each needs something usually unavailable: the training corpus, a well-chosen test statistic, or foresight at dataset release.
By Florian Braun
arXiv:2609.38917v1 Announce Type: new
Abstract: A classifier's conditional accuracy can change while its confidence distribution stays exactly the same. We study the worst-case movement of the reliab...
By Wenhao Liang, Lin Yue, Wei Emma Zhang, Mingyu Guo, Olaf Maennel, Weitong Chen
arXiv:2608.29705v1 Announce Type: cross
Abstract: Feed-forward 3D reconstruction models emit a per-pixel confidence that downstream systems read as a reliability signal. It is trained as a loss weigh...
By Nanxing Nick Deng, Qing Cheng, Niclas Zeller, Daniel Cremers
The paper audits the confidence outputs of seven feed‑forward 3D reconstruction backbones across 13 datasets, evaluating four properties: error ranking, average error‑to‑uncertainty ratio, slope of this ratio, and coverage of the implied error distribution. While confidence ranks errors well, the decoded uncertainty is consistently too small—off by at least 2.4× on median cases—and worsens with higher confidence. A post‑hoc power‑law fit per backbone‑dataset pair improves all four metrics at the dataset level, reducing the median error by 1.35×, but fails to correct coverage for many held‑out scenes, indicating the models lack the correct error scale and distribution shape.
By Nanxing Nick Deng, Qing Cheng, Niclas Zeller, Daniel Cremers
arXiv:2606. 29484v1 Announce Type: cross Abstract: Modern deepfake detectors are rarely consumed as bare classifiers.
By Md Anas Biswas
arXiv:2607. 08065v1 Announce Type: new Abstract: LLM-as-judge (Zheng et al.
By Kaihua Ding