The paper introduces CLEAR, a CLoze-style rEAsoning-based Re-ranking framework for Compositional Zero-Shot Learning. CLEAR treats primitive variations as context-driven activations of concrete visual cues rather than independent entities, extracting conditional variants in a coarse-to-fine manner and performing cloze-style reasoning to infer high-level semantics. Experiments show that CLEAR consistently improves base models and surpasses state-of-the-art methods on the C-GQA and MIT-States datasets.
By Weize Li, Zhicheng Zhao, Fei Su
The paper introduces PromptCCZSL, a framework that enables vision‑language models to continually learn new attributes, objects, and their unique compositions while avoiding forgetting. It uses a frozen VLM backbone with prompt‑based techniques, recency‑weighted multi‑teacher distillation, and several loss functions (CAL, OPL, IDL) to maintain prior knowledge and promote diverse, distinct embeddings. Experiments on UT‑Zappos and C‑GQA show significant performance gains over existing VLM‑based and non‑VLM baselines, establishing a new benchmark for continual compositional zero‑shot learning.
By Sauda Maryam, Sara Nadeem, Faisal Qureshi, Mohsen Ali
arXiv:2602. 14344v2 Announce Type: replace-cross Abstract: We study instruction following in multi-task reinforcement learning, where an agent must zero-shot execute novel tasks not seen during training.
By Mathias Jackermeier, Mattia Giuri, Jacques Cloete, Alessandro Abate
arXiv:2606. 31222v1 Announce Type: new Abstract: Composed image retrieval requires identifying a target image from a gallery by integrating a reference image with a textual modification instruction.
By Gunho Jung, Jeong-Woo Park, Seon Bin Kim, Seong-Whan Lee
ReHoPER is an inference‑only, zero‑shot method that enhances large language models’ reasoning by generating and answering intermediate questions along multiple paths before producing a final answer. It plans a horizon of candidate intermediate questions, selects one to answer, and replans based on the updated history. The approach is task‑agnostic, using generic instructions across datasets and models without labeled data or task‑specific prompt design, and it outperforms strong baselines on several datasets, notably achieving the largest gains on the new iLLC benchmark for compositional reasoning.
By Saeed Ahmadnia, Cornelia Caragea
arXiv:2604. 09686v2 Announce Type: replace Abstract: Traditional neural network models for intent inference rely heavily on observable states and struggle to generalize across diverse tasks and dynamic environments.
By Anshul Nayak, Shahil Shaik, Yue Wang