How to make effective use of domain experts for image classification?
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
Vision-language models (VLMs) such as CLIP enable zero-shot classification by comparing image features with text prompts in a shared embedding space. A fundamental property underlying this capability is the global comparability of logits across arbitrary candidate classes.
CRISP (Compositional Relational Invariance from Spatial Primitives) is an image‑classification framework that decomposes visual recognition into primitive elements and their relational composition. It represents these compositions with soft unary, binary, and ternary predicates over primitive locations and appearance, enabling differentiable spatial and visual alignment learned end‑to‑end. Evaluated on five DomainBed datasets covering style, provenance, and camera‑trap shifts, CRISP achieves new state‑of‑the‑art performance on both benchmarks.
The paper introduces a meta‑learning framework that uses a rich set of dataset‑complexity meta‑features to predict the accuracy of different classifiers on image datasets, avoiding exhaustive training. By extracting features with autoencoders, pre‑trained networks, and dimensionality reduction, regression models estimate classifier accuracies, while clustering groups similar performers to simplify recommendations. Tested on 56 diverse image datasets, the method achieves over 86% ranking prediction accuracy, offering a scalable, interpretable solution for model selection and cost reduction.
arXiv:2607. 18695v1 Announce Type: cross Abstract: A popular route to interpretable zero-shot classification asks a large language model (LLM) to describe each class name and prompts CLIP with the resulting descriptors.
arXiv:2512. 15748v2 Announce Type: replace Abstract: Visual Species Recognition (VSR) is a fundamental task in scientific disciplines that require species-level identification, including ecology, palynology, evolutionary biology, systematics, and phylogenetics.
arXiv:2609.38603v1 Announce Type: new Abstract: While earth observation models have advanced substantially, they still lack interpretability. While concept-bottleneck models provide interpretability...