arXiv Machine Learning By Aditi Naiknaware, Salimeh Sekeh

T-QPM: Enabling Temporal Out-Of-Distribution Detection and Domain Generalization for Vision-Language Models in Open-World

Read the original on arXiv Machine Learning →

arXiv:2603. 18481v2 Announce Type: replace-cross Abstract: Out-of-distribution (OOD) detection remains a critical challenge in open-world learning, where models must adapt to evolving data distributions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
2d ago

ProtoDCS: Towards Robust and Efficient Open-Set Test-Time Adaptation for Vision-Language Models

ProtoDCS introduces a robust open‑set test‑time adaptation framework for vision‑language models, addressing the challenge of simultaneously handling covariate‑shifted in‑distribution (csID) and out‑of‑distribution (csOOD) data. It replaces brittle thresholding with a double‑check separation using a probabilistic Gaussian Mixture Model and employs an evidence‑driven adaptation strategy that updates prototypes efficiently, reducing overconfidence and computational cost. Experiments on CIFAR‑10/100‑C and Tiny‑ImageNet‑C show state‑of‑the‑art performance, improving both known‑class accuracy and OOD detection metrics.

By Wei Luo, Yangfan Ou, Jin Deng, Zeshuai Deng, Xiquan Yan, Zhiquan Wen, Mingkui Tan
arXiv AI
Aug 25

What's the Catch? Evaluating Temporal Consistency in Vision-Language Models

The paper introduces TimeCatch, a benchmark that evaluates temporal consistency in vision‑language models (VLMs) by treating temporal grounding as an anomaly detection problem. Temporal anomalies are created by swapping consecutive frames, while frame‑level anomalies involve replacing a frame with Gaussian noise. Across synthetic and real‑world datasets, VLMs reliably detect and localize frame‑level anomalies but perform near chance on temporal anomaly detection, whereas humans excel at both tasks.

By Marek Hradil, Danae S\'anchez Villegas
arXiv AI
Aug 20

From Inference to Adaptation: A Unified Optimal Transport View of Vision Language Model

arXiv:2608. 18339v1 Announce Type: cross Abstract: Vision-language models (VLMs) have demonstrated remarkable zero-shot capabilities yet remain sensitive to real-world distribution shifts during inference.

By Qi Yu, Zhichen Zeng, Katherine Tieu, Xiyuan Yang, Ruizhong Qiu, Yuchen Yan, Lihui Liu, Yanjun Zhao, Lingjie Chen, Jingrui He, Hanghang Tong