Self-Improving Diffusion Classifiers with Minority Preference Optimization
arXiv:2607. 03770v1 Announce Type: cross Abstract: Prior studies have demonstrated that diffusion classifiers achieve robust zero-shot classification performance.
arXiv:2607. 03831v1 Announce Type: cross Abstract: Diffusion models have recently been repurposed for zero-shot classification, giving rise to diffusion classifiers that identify the best-matching text prompt by minimizing the noise-prediction error.
arXiv:2607. 03770v1 Announce Type: cross Abstract: Prior studies have demonstrated that diffusion classifiers achieve robust zero-shot classification performance.
The paper presents the first systematic reliability evaluation of diffusion-based Large Vision‑Language Models (dLVLMs), comparing six diffusion models to autoregressive (AR) baselines across four dimensions. Key findings include a reversal of the yes‑bias seen in AR models for binary visual queries, competitive hallucination rates but lower linguistic quality, near‑zero accuracy for underrepresented racial groups with opposite‑polarity gender bias, and accuracy collapse in multiple‑choice tasks when the correct option is shorter than distractors due to a length prior emerging at the first denoising step. Additionally, tokens committed late in denoising with low confidence correlate with hallucinated content, indicating a unique mechanistic signal in diffusion generation.
arXiv:2608. 14649v1 Announce Type: new Abstract: We present dLLM-SetScore, a training-free method that uses discrete masked-diffusion language models for multi-label text classification.
arXiv:2607. 12464v1 Announce Type: cross Abstract: When labeled data are scarce, off-the-shelf diffusion models can augment training sets for few-shot medical image classification, but not all generated samples are equally useful for the downstream task.
arXiv:2608. 19871v1 Announce Type: new Abstract: Compositional Zero-Shot Learning (CZSL) aims to recognize unseen attribute-object compositions by leveraging knowledge of primitive concepts learned from seen compositions.
arXiv:2609.37638v1 Announce Type: new Abstract: Current explanation methods for contrastive vision--language models such as CLIP mainly identify important regions without showing how to change the in...
arXiv:2512. 08724v3 Announce Type: replace Abstract: Text-to-image (TTI) diffusion models have achieved remarkable visual quality, yet they have been repeatedly shown to exhibit social biases across sensitive attributes such as gender, race and age.
Compositional Zero-Shot Learning (CZSL) aims to recognize unseen attribute-object compositions by leveraging knowledge of primitive concepts learned from seen compositions. Although recent works achieve impressive performance in CZSL by leveraging large vision-language models, they primarily rely on discriminative representations that may not explicitly preserve the structured relationships between primitive concepts and their compositions.
The paper introduces FailSAE, a method that uses Sparse Autoencoders to predict failures in vision‑language models (VLMs) such as CLIP. By treating failure prediction as a classification over sparse SAE latent activations and employing a three‑stage training pipeline, the approach yields higher prediction accuracy than existing confidence‑score or auxiliary‑classifier baselines. Analysis shows that the SAE captures class‑specific concepts and reveals a shift toward ambiguous or style‑related concepts during failures, offering insights for runtime failure recovery.
arXiv:2602. 06806v2 Announce Type: replace-cross Abstract: Text-to-image diffusion models achieve impressive generation quality but inherit and amplify training-data biases, skewing coverage of semantic attributes.
arXiv:2607. 22544v1 Announce Type: new Abstract: Visual counterfactual explanations aim to answer "what minimal change to this image would flip the model's prediction?
arXiv:2607. 18695v1 Announce Type: cross Abstract: A popular route to interpretable zero-shot classification asks a large language model (LLM) to describe each class name and prompts CLIP with the resulting descriptors.