arXiv Machine Learning

A Spectral Phase Diagram for Binary Few-Shot Classification: Intrinsic Dimensionality, Geometric Saturation, and Representational Diagnosis

arXiv:2606. 24903v1 Announce Type: new Abstract: Deciding when to stop collecting labeled examples is a fundamental but undertheorized problem in applied machine learning.

arXiv Machine Learning
Jul 14

The Geometry of Saturation: Effective Rank Predicts When Labels Stop Helping in Few-Shot Classification

arXiv:2606. 24903v2 Announce Type: replace Abstract: Few-shot label acquisition lacks a label-free signal for when additional labels cease to improve accuracy: existing stopping criteria either require a held-out validation set (violating the few-shot premise) or rely on theoretically ungrounded heuristics, so we introduce the spectral saturation index $S(K)=\mathrm{erank}(\hat{\Sigma}_W^{(K)})/K$, the exponential spectral entropy of the pooled within-class covariance normalized by per-class support size $K$, which measures the exploration rate per label and falls below a fixed threshold $\tau=0.

By Arnav Gupta
arXiv Machine Learning
1d ago

How Many Categories Are Enough? Distribution-Free Certification Limits for Few-Shot Anomaly Thresholds

The paper investigates how many normal samples are required to reliably set an alarm threshold for few‑shot anomaly detectors, focusing on distribution‑free certification limits. Using a frozen DINOv2 PCA residual ranker on 15 MVTec and 12 VisA categories, the authors show that simple leave‑one‑image‑out calibration is limited by resolution and shift, leading to empirical false‑alarm rates far above the nominal level. They derive a category‑count feasibility calculus, demonstrating that at least 14, 29, and 59 independent category draws are needed for 95% upper confidence bounds at α=0.20, 0.10, and 0.05, and propose the CRESS protocol to split source categories into reference, proposal, and certification roles. whyItMatters":"The study provides concrete numerical thresholds for the amount of source evidence needed to guarantee reliable anomaly detection in new categories, informing practical deployment of few‑shot detectors."

By Gia Huy Thai, Nguyen Thai Anh
arXiv Statistics ML
Aug 24

EDGE: a closed-form directed test for the calibration of probabilistic binary classifiers

The paper introduces EDGE, a closed‑form statistical test for assessing the calibration of probabilistic binary classifiers, specifically logistic regression. EDGE uses the same binned predicted‑versus‑observed table as a reliability diagram, projects standardized bin residuals onto a small basis of smooth calibration‑distortion shapes, and yields a null distribution that is a weighted sum of chi‑square variables. The method requires only a single pass over the data and a small eigendecomposition, avoiding refitting, resampling, or tuning, and remains robust in sparse or misspecified settings where other binned tests fail.

By Ebrahim Khaled Ebrahim, Ahmed El-Kotory
arXiv AI
Sep 7

Phase Transition Frequency as a Training Time Predictor of Test Accuracy in ResNets

The study investigates whether the number of discrete class‑separability jumps (phase transitions) observed during ResNet fine‑tuning can predict final test accuracy. Across 75 experiments on four benchmarks (CIFAR‑10, CIFAR‑100, TinyImageNet, CIFAR‑10‑C) and three ResNet variants, a strong negative correlation is found on standard i.i.d. datasets (r = −0.84 on CIFAR‑10, r = −0.87 on CIFAR‑100), while the correlation weakens under distributional stress. Additional analyses show that the transition count retains predictive power after controlling for architecture depth and outperforms other training‑curve signals on in‑distribution benchmarks, though it is dominated by other signals on stressed datasets.

By Arunan J
arXiv AI
Sep 11

OmniMed-FL: A Robust Multimodal Federated Learning Framework for Clinical Diagnosis

OmniMed‑FL is a multimodal federated learning framework that fuses chest radiographs and synthetic patient notes to classify five clinical conditions. The study benchmarks eight fusion strategies, three initializations, and four missing‑text imputation rules across 3–20 hospital clients under non‑IID Dirichlet partitioning, showing that federated approaches (FedAvg, FedProx, SCAFFOLD‑AdamW) outperform local‑only training. Multimodal fusion consistently improves performance, achieving macro‑F1 scores up to 0.956 on the synthetic corpus and 0.906 on the radiograph corpus.

By Ayush Debnath, Ruelia Saha, Sudip Misra
arXiv Machine Learning
Sep 25

How Many Humans Are 32 LLM Judges Worth?

The paper investigates how many human annotators are equivalent to a panel of 32 large‑language‑model (LLM) judges. By comparing the panel’s label distributions to empirical human labels on three ChaosNLI tasks, the authors find two distinct effective panel sizes: distribution‑error matching yields effective sizes of 2.304, 3.750, and 3.445, while spectral matching gives 4.242, 6.459, and 6.499, indicating a 1.72–1.89× gap. The study also explores how spectral diversity, participation ratio, and panel composition affect effective size, and demonstrates that carefully chosen panels can outperform baseline accuracy while improving effective size.

By Chao Li, Yingying Yu, Yunfeng Li
arXiv Machine Learning
Jul 31

DS@GT ARC at ImageCLEFmedical 2026: Architectural Diversity for Concept Detection and Foundation-Model Scaling for Caption Prediction in Medical Image Analysis

arXiv:2607. 27763v1 Announce Type: cross Abstract: We describe the DS@GT submissions to the ImageCLEFmedical Caption 2026 challenge, which continues a long-running benchmark on the ROCOv2 dataset with two tracks: Concept Detection (Task 1), assigning UMLS Concept Unique Identifiers (CUIs) to radiology images, and Caption Prediction (Task 2), generating natural-language captions.

By Bowen Wang, Youwen Zhang, Ritesh Mehta