arXiv Computer Vision By Weize Li, Zhicheng Zhao, Fei Su

From Model Patterns to Abstract Semantics in Compositional Zero-Shot Learning

Read the original on arXiv Computer Vision →

The paper introduces CLEAR, a CLoze-style rEAsoning-based Re-ranking framework for Compositional Zero-Shot Learning. CLEAR treats primitive variations as context-driven activations of concrete visual cues rather than independent entities, extracting conditional variants in a coarse-to-fine manner and performing cloze-style reasoning to infer high-level semantics. Experiments show that CLEAR consistently improves base models and surpasses state-of-the-art methods on the C-GQA and MIT-States datasets.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

Hugging Face Trending Papers
Aug 20

DIFFCZSL: Compositional Zero-Shot Learning Regularized by Diffusion Representations

Compositional Zero-Shot Learning (CZSL) aims to recognize unseen attribute-object compositions by leveraging knowledge of primitive concepts learned from seen compositions. Although recent works achieve impressive performance in CZSL by leveraging large vision-language models, they primarily rely on discriminative representations that may not explicitly preserve the structured relationships between primitive concepts and their compositions.

arXiv AI
Sep 10

A Progressive Training Strategy for Embodied Vision-Language Models to Mitigate Spatio-Temporal Hallucinations

The paper introduces a progressive training strategy for embodied vision‑language models aimed at reducing spatio‑temporal hallucinations. It first creates a Chain‑of‑Thought dataset that breaks complex reasoning into detailed spatiotemporal steps, then uses supervised pre‑training on this dataset followed by fine‑tuning with weakly‑labeled data. Experiments show the method improves backbone accuracy and narrows the forward‑backward performance gap from over 70% to 6.53%, indicating stronger dynamic reasoning and fewer temporal biases.

By Xiaoda Yang, Shuai Yang, Can Wang, Jingyang Xue, Menglan Tang, Checheng Yu, Xunzhe Zhou, Sashuai Zhou, Tao Jin, Lixin Yang, Xiangyu Yue, Zhou Zhao