arXiv:2608. 13190v1 Announce Type: new Abstract: Group-robust learning is crucial for maintaining accuracy on rare subpopulations when training-group labels are unavailable.
By Qianqian Wang, Yunshan Li, Dawei Huang, Wenwu Gong, Lili Yang
Test-time adaptation (TTA) can mitigate domain shift without source data, but it is highly brittle under adversarially contaminated test streams, where corrupted inputs also destabilize online updates. We study robust test-time adaptation (RTTA) in the adversarial-stream setting, which remains comparatively underexplored relative to standard TTA, and propose SAFER (Stochastic Augmentation Framework for Enhanced Robustness), a training-free reliability-guided augmentation wrapper for RTTA.
arXiv:2608. 09768v1 Announce Type: new Abstract: A prediction that is both confident and wrong is a critical reliability failure because it can bypass abstention and human review precisely when the model is mistaken.
By Ange-Cl\'ement Akazan, Ineza Remy Mugenga, Abebe Geletu, Jean Medard Ngnotchouye, Issa Karambal
arXiv:2608. 04442v1 Announce Type: new Abstract: Robustness to natural corruptions remains a fundamental challenge for deep neural networks.
By Jiangang Yang, Wenhui Shi, Lu Hu, Jing Xing, Jian Liu
arXiv:2609.16380v1 Announce Type: new
Abstract: Class-balanced learning and label noise create a coupled failure mode: frequency correction prevents majority classes from dominating the decision rule...
By Mushir Akhtar, Akarsh J., M. Tanveer, Mohd. Arshad
arXiv:2608. 14089v1 Announce Type: new Abstract: Safety classifiers deployed with large language models often fail for two reasons: their decisions reflect the policy learned during training rather than the deployer's desired policy, and their performance degrades as deployment traffic evolves.
By Thiago Sandoval, Ufuk Topcu
CORE-STACK+ is a new meta‑learning framework for deep stacked generalization that tackles two key problems in heterogeneous vision ensembles: prediction‑space multicollinearity and calibration collapse. It introduces a four‑step preconditioning pipeline—kernelized redundancy filtering, a lightweight differentiable meta‑feature gate, a spectrum‑adaptive ridge penalty, and a Laplace‑approximate Bayesian blender—to jointly improve conditioning and calibration. Across six vision benchmarks, CORE‑STACK+ boosts accuracy, reduces model count and inference cost, and significantly lowers expected calibration error compared to existing methods.
By Noor Islam S. Mohammad
arXiv:2606. 00320v1 Announce Type: new Abstract: We present an online, distribution-free framework for controlling the Conditional Value-at-Risk (CVaR), extending conformal tail risk control to non-stationary and adversarial environments.
By Catherine Chen, Jingyan Shen, Zhun Deng, Lihua Lei
The paper introduces MuViS-C, a multi‑domain benchmark that evaluates the robustness of learning‑based virtual sensing models against ten common sensor failure modes, ranging from subtle drifts to catastrophic dropouts. It assesses models using average error, relative degradation, and worst‑case fragility across nine datasets from six domains, comparing six architectures (gradient‑boosted trees, convolution, recurrence, attention, and MLP‑mixing). The study finds that all models degrade under corruption, gradient‑boosted trees are most robust, and targeted robustification can improve attention models at the cost of nominal performance.
By Jens U. Brandt, Noah C. Puetz, Alexander Windmann, Marc Hilbert, Elena Raponi, Thomas B\"ack, Thomas Bartz-Beielstein
Virtual sensing, the estimation of hard-to-measure quantities from available sensor measurements, is a critical enabler for control and monitoring in cyber-physical systems. However, when sensors fail...
arXiv:2607. 06637v1 Announce Type: new Abstract: In this work, we propose a unified approach for diagnosing misclassification and assessing the robustness of black-box classifiers.
By Evgenii Kuriabov, David Miller, Jia Li
The paper introduces Calibration-Aware Uncertainty Cascades (CAUC), a post‑hoc framework that calibrates each model’s confidence independently and uses these calibrated scores to decide when to accept an early prediction, invoke a stronger model, or combine outputs. CAUC establishes a common reliability scale across heterogeneous models, decoupling deployment policies from specific model pools or budgets. Experiments on six language benchmarks show a 1.9% relative accuracy gain over strong‑model‑only inference while cutting strong‑model calls by about 47%, and on image classification it maintains or improves performance while reducing GFLOPs by up to 57%.
By Yilin Zhang, Han Jiang, Cai Xu, Ying Liu, Wei Zhao