arXiv Statistics ML

DeepGOF-1: A Pretrained Convolutional Goodness-of-Fit Test for Logistic Regression with a Computable Consistency Certificate

DeepGOF-1 introduces a pretrained convolutional network as a goodness‑of‑fit test for logistic regression, where the network reads a grid of standardized residuals as an image and outputs a test statistic. The test is fully calibrated via the analyst’s own bootstrap, ensuring the nominal level is maintained regardless of the network’s training. The authors prove exactness under pivotality, asymptotic exactness without it, and provide a computable consistency certificate from the frozen weights, demonstrating superior stability and power across multiple benchmarks and sample sizes.

arXiv AI
Sep 7

Phase Transition Frequency as a Training Time Predictor of Test Accuracy in ResNets

The study investigates whether the number of discrete class‑separability jumps (phase transitions) observed during ResNet fine‑tuning can predict final test accuracy. Across 75 experiments on four benchmarks (CIFAR‑10, CIFAR‑100, TinyImageNet, CIFAR‑10‑C) and three ResNet variants, a strong negative correlation is found on standard i.i.d. datasets (r = −0.84 on CIFAR‑10, r = −0.87 on CIFAR‑100), while the correlation weakens under distributional stress. Additional analyses show that the transition count retains predictive power after controlling for architecture depth and outperforms other training‑curve signals on in‑distribution benchmarks, though it is dominated by other signals on stressed datasets.

By Arunan J
arXiv Computer Vision
Sep 11

A Calibration Audit of Confidence in Feed-Forward 3D Reconstruction Models

The paper audits the confidence outputs of seven feed‑forward 3D reconstruction backbones across 13 datasets, evaluating four properties: error ranking, average error‑to‑uncertainty ratio, slope of this ratio, and coverage of the implied error distribution. While confidence ranks errors well, the decoded uncertainty is consistently too small—off by at least 2.4× on median cases—and worsens with higher confidence. A post‑hoc power‑law fit per backbone‑dataset pair improves all four metrics at the dataset level, reducing the median error by 1.35×, but fails to correct coverage for many held‑out scenes, indicating the models lack the correct error scale and distribution shape.

By Nanxing Nick Deng, Qing Cheng, Niclas Zeller, Daniel Cremers