arXiv Machine Learning By Arnav Gupta

A Spectral Phase Diagram for Binary Few-Shot Classification: Intrinsic Dimensionality, Geometric Saturation, and Representational Diagnosis

Read the original on arXiv Machine Learning →

arXiv:2606. 24903v1 Announce Type: new Abstract: Deciding when to stop collecting labeled examples is a fundamental but undertheorized problem in applied machine learning.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jul 14

The Geometry of Saturation: Effective Rank Predicts When Labels Stop Helping in Few-Shot Classification

arXiv:2606. 24903v2 Announce Type: replace Abstract: Few-shot label acquisition lacks a label-free signal for when additional labels cease to improve accuracy: existing stopping criteria either require a held-out validation set (violating the few-shot premise) or rely on theoretically ungrounded heuristics, so we introduce the spectral saturation index $S(K)=\mathrm{erank}(\hat{\Sigma}_W^{(K)})/K$, the exponential spectral entropy of the pooled within-class covariance normalized by per-class support size $K$, which measures the exploration rate per label and falls below a fixed threshold $\tau=0.

By Arnav Gupta
arXiv Machine Learning
1d ago

How Many Categories Are Enough? Distribution-Free Certification Limits for Few-Shot Anomaly Thresholds

The paper investigates how many normal samples are required to reliably set an alarm threshold for few‑shot anomaly detectors, focusing on distribution‑free certification limits. Using a frozen DINOv2 PCA residual ranker on 15 MVTec and 12 VisA categories, the authors show that simple leave‑one‑image‑out calibration is limited by resolution and shift, leading to empirical false‑alarm rates far above the nominal level. They derive a category‑count feasibility calculus, demonstrating that at least 14, 29, and 59 independent category draws are needed for 95% upper confidence bounds at α=0.20, 0.10, and 0.05, and propose the CRESS protocol to split source categories into reference, proposal, and certification roles. whyItMatters":"The study provides concrete numerical thresholds for the amount of source evidence needed to guarantee reliable anomaly detection in new categories, informing practical deployment of few‑shot detectors."

By Gia Huy Thai, Nguyen Thai Anh
arXiv Statistics ML
Aug 24

EDGE: a closed-form directed test for the calibration of probabilistic binary classifiers

The paper introduces EDGE, a closed‑form statistical test for assessing the calibration of probabilistic binary classifiers, specifically logistic regression. EDGE uses the same binned predicted‑versus‑observed table as a reliability diagram, projects standardized bin residuals onto a small basis of smooth calibration‑distortion shapes, and yields a null distribution that is a weighted sum of chi‑square variables. The method requires only a single pass over the data and a small eigendecomposition, avoiding refitting, resampling, or tuning, and remains robust in sparse or misspecified settings where other binned tests fail.

By Ebrahim Khaled Ebrahim, Ahmed El-Kotory
arXiv AI
Sep 7

Phase Transition Frequency as a Training Time Predictor of Test Accuracy in ResNets

The study investigates whether the number of discrete class‑separability jumps (phase transitions) observed during ResNet fine‑tuning can predict final test accuracy. Across 75 experiments on four benchmarks (CIFAR‑10, CIFAR‑100, TinyImageNet, CIFAR‑10‑C) and three ResNet variants, a strong negative correlation is found on standard i.i.d. datasets (r = −0.84 on CIFAR‑10, r = −0.87 on CIFAR‑100), while the correlation weakens under distributional stress. Additional analyses show that the transition count retains predictive power after controlling for architecture depth and outperforms other training‑curve signals on in‑distribution benchmarks, though it is dominated by other signals on stressed datasets.

By Arunan J