arXiv AI

Beyond Output-Space Calibration: Spectral Evidence Bundling for Selective Reliability Estimation in Time-Series Classification

arXiv:2607. 18279v1 Announce Type: cross Abstract: Post-hoc calibration for time-series classification usually remaps output scores, but deployment decisions such as trust, abstention, and review depend on whether a confident prediction is supported by the current temporal signal.

arXiv Machine Learning
Jul 2

LeNEPA: No-Augmentation Next-Latent Prediction for Time-Series Representation Learning

arXiv:2607. 00958v1 Announce Type: new Abstract: Time series are central to modern data mining applications, from industrial telemetry and server metrics to finance and physiology, yet time-series self-supervised learning often depends on view and augmentation choices that encode domain-specific invariances.

By Alexander Chemeris, Ming Jin, Randall Balestriero
arXiv AI
Sep 2

Counterfactual Fragility Certificates: Exposing High-Confidence Brittleness under Structured Evidence Failure

The paper introduces Counterfactual Fragility Certificates (CFC), a model‑agnostic audit protocol that maps each prediction to an evidence‑failure trajectory, summarizing it with metrics such as greedy flip budget, margin‑collapse area, degradation thresholds, and fragility dominance score. CFC is shown to identify brittle high‑confidence predictions on seven tabular benchmarks with an AUROC of 0.915, outperforming existing scalar scores by up to +0.405. The method remains effective across various perturbation and review‑budget scenarios, and can also inform fragility‑aware regularization and temperature correction.

By Filippo Cenacchi, Longbing Cao, Runze Yang
arXiv Statistics ML
Aug 24

EDGE: a closed-form directed test for the calibration of probabilistic binary classifiers

The paper introduces EDGE, a closed‑form statistical test for assessing the calibration of probabilistic binary classifiers, specifically logistic regression. EDGE uses the same binned predicted‑versus‑observed table as a reliability diagram, projects standardized bin residuals onto a small basis of smooth calibration‑distortion shapes, and yields a null distribution that is a weighted sum of chi‑square variables. The method requires only a single pass over the data and a small eigendecomposition, avoiding refitting, resampling, or tuning, and remains robust in sparse or misspecified settings where other binned tests fail.

By Ebrahim Khaled Ebrahim, Ahmed El-Kotory