arXiv:2609.01027v1 Announce Type: new
Abstract: Out-of-distribution (OOD) detection predicts whether a test image belongs to none of the predefined classes. To evaluate this task, benchmarks need ima...
By Ruslan Rozumnyi, Mat\v{e}j Such\'anek, Tom\'a\v{s} Voj\'i\v{r}, Kl\'ara Janou\v{s}kov\'a, Ji\v{r}\'i Matas
ProtoDCS introduces a robust open‑set test‑time adaptation framework for vision‑language models, addressing the challenge of simultaneously handling covariate‑shifted in‑distribution (csID) and out‑of‑distribution (csOOD) data. It replaces brittle thresholding with a double‑check separation using a probabilistic Gaussian Mixture Model and employs an evidence‑driven adaptation strategy that updates prototypes efficiently, reducing overconfidence and computational cost. Experiments on CIFAR‑10/100‑C and Tiny‑ImageNet‑C show state‑of‑the‑art performance, improving both known‑class accuracy and OOD detection metrics.
By Wei Luo, Yangfan Ou, Jin Deng, Zeshuai Deng, Xiquan Yan, Zhiquan Wen, Mingkui Tan
arXiv:2608. 01074v1 Announce Type: new Abstract: Tabular data is used extensively in many real-world use cases.
By Mayank Sharma, Rohit Kumar Mourya, Pratik Mazumder
arXiv:2607.09086v2 Announce Type: replace
Abstract: We present Subtoken Vision Transformer (SubViT), a selective image tokenization method for fine-grained visual recognition. Standard Vision Transfo...
By Jie Zhu, Ivy Zhang, Minchul Kim, Xiaoming Liu
arXiv:2608. 03557v1 Announce Type: cross Abstract: Tabular-to-image methods that convert tabular data into visual representations have emerged as a novel paradigm for leveraging the high performance of deep learning models.
By Malena Loza, Felipe Grijalva, Eva Milara, Luis Bote-Curiel, Francisco J. Lara-Abelenda, David Chushig-Muzo
The paper introduces UnInfo, a test‑time adaptation method for vision‑language models like CLIP that addresses image corruption—a realistic distribution shift caused by sensor conditions. UnInfo leverages uniformity‑aware confidence maximization, information‑aware loss balancing, and knowledge distillation from an EMA teacher to preserve embedding uniformity and improve zero‑shot classification accuracy. Experiments show that UnInfo outperforms existing TTA methods on corrupted image datasets.
By Kazuki Adachi, Shin'ya Yamaguchi, Tomoki Hamagami