The paper introduces CLEAR, a CLoze-style rEAsoning-based Re-ranking framework for Compositional Zero-Shot Learning. CLEAR treats primitive variations as context-driven activations of concrete visual cues rather than independent entities, extracting conditional variants in a coarse-to-fine manner and performing cloze-style reasoning to infer high-level semantics. Experiments show that CLEAR consistently improves base models and surpasses state-of-the-art methods on the C-GQA and MIT-States datasets.
By Weize Li, Zhicheng Zhao, Fei Su
The paper introduces PromptCCZSL, a framework that enables vision‑language models to continually learn new attributes, objects, and their unique compositions while avoiding forgetting. It uses a frozen VLM backbone with prompt‑based techniques, recency‑weighted multi‑teacher distillation, and several loss functions (CAL, OPL, IDL) to maintain prior knowledge and promote diverse, distinct embeddings. Experiments on UT‑Zappos and C‑GQA show significant performance gains over existing VLM‑based and non‑VLM baselines, establishing a new benchmark for continual compositional zero‑shot learning.
By Sauda Maryam, Sara Nadeem, Faisal Qureshi, Mohsen Ali
arXiv:2602. 14344v2 Announce Type: replace-cross Abstract: We study instruction following in multi-task reinforcement learning, where an agent must zero-shot execute novel tasks not seen during training.
By Mathias Jackermeier, Mattia Giuri, Jacques Cloete, Alessandro Abate
arXiv:2606. 31222v1 Announce Type: new Abstract: Composed image retrieval requires identifying a target image from a gallery by integrating a reference image with a textual modification instruction.
By Gunho Jung, Jeong-Woo Park, Seon Bin Kim, Seong-Whan Lee
ReHoPER is an inference‑only, zero‑shot method that enhances large language models’ reasoning by generating and answering intermediate questions along multiple paths before producing a final answer. It plans a horizon of candidate intermediate questions, selects one to answer, and replans based on the updated history. The approach is task‑agnostic, using generic instructions across datasets and models without labeled data or task‑specific prompt design, and it outperforms strong baselines on several datasets, notably achieving the largest gains on the new iLLC benchmark for compositional reasoning.
By Saeed Ahmadnia, Cornelia Caragea
arXiv:2604. 09686v2 Announce Type: replace Abstract: Traditional neural network models for intent inference rely heavily on observable states and struggle to generalize across diverse tasks and dynamic environments.
By Anshul Nayak, Shahil Shaik, Yue Wang
Compositional Zero-Shot Learning (CZSL) aims to recognize unseen attribute-object compositions by leveraging knowledge of primitive concepts learned from seen compositions. Although recent works achieve impressive performance in CZSL by leveraging large vision-language models, they primarily rely on discriminative representations that may not explicitly preserve the structured relationships between primitive concepts and their compositions.
The paper investigates how large language models activate a latent symbolic reasoning circuit—comprising abstraction, induction, and retrieval—when presented with in-context examples. By tracking this circuit across different shot counts and model families, the authors show that its components become detectable and functional long before the model reaches high accuracy. They further demonstrate that per-head causal contributions can increase eightfold from 1- to 10-shot, and that interventions such as cross-shot activation patching or function vector injection can dramatically improve accuracy, even at 0-shot, by leveraging the pre‑existing circuit in the model weights.
By Melissa Wessel
arXiv:2607. 00374v1 Announce Type: cross Abstract: Composed Image Retrieval (CIR) retrieves a target image from a reference image and a textual modification.
By Jingjing Zhang, Lei Zhang, Zheren Fu, Zhendong Mao
The paper introduces a dependency‑graph framework to formalize compositional reasoning in language models, defining three increasing levels of compositionality. Using data‑structure tasks with deterministic rewards, the authors observe a consistent asymmetry: training on decomposed skills does not reliably transfer to composed tasks, whereas training on composed tasks transfers more readily to decomposed ones. They provide a theoretical explanation for this asymmetry and evaluate its effects under length extrapolation, structural distribution shift, and transfer to unseen skills, concluding with a pilot study on real‑world tool‑calling benchmarks that suggests the phenomenon extends to practical settings.
By Yu He, Yingxi Li, Yifei Wang, Ellen Vitercik
arXiv:2608. 20161v1 Announce Type: new Abstract: Instruction-based image editing uses a planner-renderer pipeline: a vision-language model (VLM) first converts the instruction into an edit plan, and a diffusion model then executes that plan.
By Haoxiang Cao, Jiajiong Cao, Xuanpu Zhang, Changqian Yu, Chaoqun Wang
arXiv:2609.35942v1 Announce Type: new
Abstract: Recent work in visual question answering has shown that vision-language models can exhibit strong reasoning capabilities by translating visual inputs i...
By Ting-Chih Chen, Emile van Krieken, Shujian Yu, Filip Ilievski