arXiv:2506. 14194v2 Announce Type: replace Abstract: We present a theory for the construction of out-of-distribution (OOD) detection features for neural networks.
By Sudeepta Mondal (Mary), Xinyi (Mary), Xie, Alex Wong, Ganesh Sundaramoorthi
arXiv:2503. 05169v2 Announce Type: replace Abstract: Applying machine learning to increasingly high-dimensional problems with sparse or biased training data increases the risk that a model is used on inputs outside its training domain.
By Felix Krumbiegel, Juniper Tyree, Michael Boy, Petri Clusius, Andreas Rupp
arXiv:2606. 16196v1 Announce Type: new Abstract: Deep neural networks have achieved remarkable performance across medical imaging tasks, yet their tendency to overgeneralize under distributional shifts poses a major obstacle to safe clinical deployment.
By Anju Chhetri, Pratik Shrestha, Ramesh Rana, Prashnna Gyawali, Binod Bhattarai
Detecting out-of-distribution (OOD) data is crucial for reliable machine learning deployment. Among detection strategies, post-hoc methods are particularly attractive due to their efficiency, as they operate directly on pre-trained networks without requiring retraining.
The paper introduces a framework for out-of-distribution (OOD) detection that addresses the trade‑off between detection performance and classification accuracy caused by fine‑tuning with auxiliary outlier data. It optimizes three factors—model reminder, data sampling, and representation learning—by proposing Self‑Knowledge Distillation to preserve accuracy, Semi‑hard Outlier Sampling to enhance detection with minimal data, and Outlier‑aware Supervised Contrastive Learning to improve ID‑OOD separability. The combined approach yields cumulative gains, outperforming existing methods on diverse benchmarks, especially in long‑tailed scenarios, and offers a robust baseline for real‑world OOD detection.
By Hyunjun Choi, JaeHo Chung, Hawook Jeong
arXiv:2604. 08572v2 Announce Type: replace Abstract: State-of-the-art post-hoc out-of-distribution detection methods rely on intermediate layer activation editing.
By Gianluca Guglielmo, Marc Masana
arXiv:2508.10148v2 Announce Type: replace-cross
Abstract: Accurate and explainable out-of-distribution (OOD) detection is required to use machine learning systems safely. Previous work has shown that...
By Maria Stoica, Francesco Leofante, Alessio Lomuscio
The paper introduces Dynamic DAE Guardrails (DSG), a method that uses Dynamic Sparse Autoencoders to perform precision unlearning in large language models. DSG leverages principled feature selection and a dynamic classifier to target activation-based unlearning, outperforming existing gradient‑based methods in terms of computational efficiency, stability, sequential unlearning, resistance to relearning attacks, data efficiency, and interpretability.
By Aashiq Muhamed, Jacopo Bonato, Mona Diab, Virginia Smith
arXiv:2606. 12138v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) are widely used to interpret neural network representations, but their utility depends on whether the learned features are reproducible across training runs.
By Gleb Gerasimov, Timofei Rusalev, Nikita Balagansky, Daniil Laptev, Vadim Kurochkin, Daniil Gavrilov
arXiv:2606. 29952v1 Announce Type: cross Abstract: Detecting out-of-distribution (OOD) data is crucial for reliable machine learning deployment.
By Seonghwan Park, Hyunji Jung, Dongyeop Lee, Namhoon Lee
arXiv:2605. 28021v2 Announce Type: replace Abstract: Out-of-distribution (OOD) detection is essential for deploying machine learning models in open-world and safety-critical scenarios, where test inputs may deviate from the training distribution and overconfident predictions on unknown samples can lead to unreliable decisions.
By Fengqiang Wan, Qing-Yuan Jiang, Fu Shen, Yang Yang
arXiv:2409. 10094v3 Announce Type: replace-cross Abstract: Out-of-Distribution (OoD) detection aims to justify whether a given sample is from the training distribution of the classifier-under-protection, i.
By Kun Fang, Zuopeng Yang, Haibo Hu, Xiaolin Huang, Jie Yang, Qinghua Tao