arXiv AI By Xizhan Gao, Wei Hu

DCSCR: A Class-Specific Collaborative Representation based Network for Image Set Classification

Read the original on arXiv AI →

arXiv:2508. 12745v2 Announce Type: replace-cross Abstract: Image set classification (ISC), which can be viewed as a task of comparing similarities between sets consisting of unordered heterogeneous images with variable quantities and qualities, has attracted growing research attention in recent years.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Sep 24

CRISP: Compositional Relations as Invariant Structural Priors for Domain Generalization

CRISP (Compositional Relational Invariance from Spatial Primitives) is an image‑classification framework that decomposes visual recognition into primitive elements and their relational composition. It represents these compositions with soft unary, binary, and ternary predicates over primitive locations and appearance, enabling differentiable spatial and visual alignment learned end‑to‑end. Evaluated on five DomainBed datasets covering style, provenance, and camera‑trap shifts, CRISP achieves new state‑of‑the‑art performance on both benchmarks.

By Dat Nguyen, Duc-Duy Nguyen
Hugging Face Trending Papers
Jun 3

COMBINER: Composed Image Retrieval Guided by Attribute-based Neighbor Relations

Composed Image Retrieval (CIR) represents a challenging retrieval task that targets locating specific images through multimodal inputs. Despite recent progress in CIR techniques, prior approaches often overlook cases where images appear visually alike yet differ in attributes, potentially undermining both multimodal feature fusion and similarity modeling.

arXiv Computer Vision
Sep 7

Few-Shot Video Recognition via Hierarchical Metric Learning

The paper introduces HML-FSAR, a hierarchical metric learning framework for few-shot action recognition. It incorporates a spatial‑enhanced module, temporal MHA, heterogeneous alignment, spatial‑temporal fusion, and dictionary learning to build a comprehensive feature pipeline. Progressive constraints—center, alignment, contrastive, dictionary, and prototype metrics—are applied from frame‑level representations to final prototypes, improving feature compactness, alignment, discriminability, and robustness.

By Jiaxin Zhang, Haoran Gao, Xizhan Gao, Zihao Dong, Tingwei Wang, Sijie Niu
Hugging Face Trending Papers
Jun 23

Training-free Cross-domain Few-shot Segmentation via Robust Semantic Representation and Matching

Cross-domain Few-shot Segmentation (CD-FSS) aims to transfer knowledge learned from source domain to distinct target domains, segmenting unseen target classes with only a few annotated samples. Although existing methods have made significant progress, they still rely on training or fine-tuning processes, which incur high computational costs and risk overfitting.