CRISP (Compositional Relational Invariance from Spatial Primitives) is an image‑classification framework that decomposes visual recognition into primitive elements and their relational composition. It represents these compositions with soft unary, binary, and ternary predicates over primitive locations and appearance, enabling differentiable spatial and visual alignment learned end‑to‑end. Evaluated on five DomainBed datasets covering style, provenance, and camera‑trap shifts, CRISP achieves new state‑of‑the‑art performance on both benchmarks.
By Dat Nguyen, Duc-Duy Nguyen
Composed Image Retrieval (CIR) represents a challenging retrieval task that targets locating specific images through multimodal inputs. Despite recent progress in CIR techniques, prior approaches often overlook cases where images appear visually alike yet differ in attributes, potentially undermining both multimodal feature fusion and similarity modeling.
The paper introduces HML-FSAR, a hierarchical metric learning framework for few-shot action recognition. It incorporates a spatial‑enhanced module, temporal MHA, heterogeneous alignment, spatial‑temporal fusion, and dictionary learning to build a comprehensive feature pipeline. Progressive constraints—center, alignment, contrastive, dictionary, and prototype metrics—are applied from frame‑level representations to final prototypes, improving feature compactness, alignment, discriminability, and robustness.
By Jiaxin Zhang, Haoran Gao, Xizhan Gao, Zihao Dong, Tingwei Wang, Sijie Niu
arXiv:2608.31107v1 Announce Type: new
Abstract: The advent of foundation models have enabled a new era in zero-shot classification. Yet, key challenges persist. Despite their impressive generalizatio...
By Lucas Wojcik, Gabriel E. Lima, Sergio M. Silva Jr., Eduil Nascimento Jr., David Menotti
Cross-domain Few-shot Segmentation (CD-FSS) aims to transfer knowledge learned from source domain to distinct target domains, segmenting unseen target classes with only a few annotated samples. Although existing methods have made significant progress, they still rely on training or fine-tuning processes, which incur high computational costs and risk overfitting.
arXiv:2607. 09104v1 Announce Type: cross Abstract: While the growing availability of image data has driven significant advances, labeling datasets remains costly and time-consuming.
By Camila Piscioneri Magalh\~aes, Lucas Pascotti Valem