arXiv:2607. 22797v1 Announce Type: cross Abstract: Trustworthy deployment of AI-based diagnosis in safety-critical mechanical systems hinges on validation: whether a prediction can be checked against physical reality before it is acted upon.
By Yuntong Chen, Jianyu Liu, Guobin Zhao, Ziang Wang, Chao Chen, Ju Huang, Xitian Tian, Lijiang Huang
arXiv:2607. 00958v1 Announce Type: new Abstract: Time series are central to modern data mining applications, from industrial telemetry and server metrics to finance and physiology, yet time-series self-supervised learning often depends on view and augmentation choices that encode domain-specific invariances.
By Alexander Chemeris, Ming Jin, Randall Balestriero
arXiv:2608. 10406v1 Announce Type: cross Abstract: Web search, product search, and question-answering retrieval systems often assign a relevance label and confidence score to each query-candidate pair.
By Inwoo Tae, Yongjae Lee
The paper introduces Counterfactual Fragility Certificates (CFC), a model‑agnostic audit protocol that maps each prediction to an evidence‑failure trajectory, summarizing it with metrics such as greedy flip budget, margin‑collapse area, degradation thresholds, and fragility dominance score. CFC is shown to identify brittle high‑confidence predictions on seven tabular benchmarks with an AUROC of 0.915, outperforming existing scalar scores by up to +0.405. The method remains effective across various perturbation and review‑budget scenarios, and can also inform fragility‑aware regularization and temperature correction.
By Filippo Cenacchi, Longbing Cao, Runze Yang
arXiv:2608.22233v1 Announce Type: cross
Abstract: Test-time adaptation (TTA) aims to improve model robustness under distribution shift by adapting a source model using unlabeled test data. Although m...
By Sreeja Guha Majumdar, Aratrika Saha
arXiv:2606. 08037v1 Announce Type: cross Abstract: Electrocardiogram (ECG) classification models often suffer from severe label scarcity, making semi-supervised learning (SSL) an attractive strategy for reducing annotation costs.
By Hongkyu Koh, Ikbeom Jang
The paper introduces EDGE, a closed‑form statistical test for assessing the calibration of probabilistic binary classifiers, specifically logistic regression. EDGE uses the same binned predicted‑versus‑observed table as a reliability diagram, projects standardized bin residuals onto a small basis of smooth calibration‑distortion shapes, and yields a null distribution that is a weighted sum of chi‑square variables. The method requires only a single pass over the data and a small eigendecomposition, avoiding refitting, resampling, or tuning, and remains robust in sparse or misspecified settings where other binned tests fail.
By Ebrahim Khaled Ebrahim, Ahmed El-Kotory
arXiv:2608. 15951v1 Announce Type: new Abstract: We study when a wearable stress system should surface a prediction rather than change it.
By Jaden Moon, Yu Wu, Arvind Pillai, Andrew Campbell
arXiv:2609.35929v1 Announce Type: new
Abstract: Time series classification (TSC) exhibits a sharp trade-off between accuracy and computational scalability. Meta-ensembles like HIVE-COTE 2.0 reach sta...
By Onisa Mpaunda
arXiv:2606. 03631v1 Announce Type: cross Abstract: Multivariate time series classification (MTSC) is pivotal in high-stakes domains, such as clinical diagnosis and industrial fault detection, where safe deployment necessitates transparent decision-making.
By Tao Xie, Zexi Tan, Haoyi Xiao, Mengke Li, Yiqun Zhang, Yang Lu, Cuie Yang, Yiu-ming Cheung
arXiv:2607. 18278v1 Announce Type: cross Abstract: Calibration is usually evaluated in aggregate, but the most dangerous failures are often local: predictions that remain highly confident despite being wrong.
By Filippo Cenacchi, Longbing Cao, Runze Yang
arXiv:2607. 20590v1 Announce Type: new Abstract: Neural network misclassifications exhibit characteristic spectral instability in internal activations that is invisible at the output layer.
By Arunan J