Hugging Face Trending Papers

Learning visual representations for compositional analysis of artworks and photographs

Composition, the deliberate arrangement of visual elements, is central to how meaning, emotion, and aesthetic quality are conveyed in artwork, yet it remains among the least formalized dimensions of visual understanding. Prior work highlights a persistent gap in learning meaningful compositional representations, attributing it to semantic bias and suggesting that human-inspired approaches may be key.

arXiv Machine Learning
Aug 28

How AI Experiences Art: Emergent Aesthetic Structure in a Self-Supervised Multimodal Embedding Space

The paper introduces a self‑supervised framework that maps text, audio, image, and video into a shared 256‑dimensional embedding space and uses iterative clustering to uncover aesthetic structure. It examines how AI’s cluster assignments diverge from human affective labels on a weakly supervised multimodal dataset. The study highlights implications for cross‑modal similarity, media organization for Retrieval‑Augmented Generation, and automated data labeling.

By Corey D. C. Heath
Hugging Face Trending Papers
Aug 27

How AI Experiences Art: Emergent Aesthetic Structure in a Self-Supervised Multimodal Embedding Space

The paper explores how AI can develop its own aesthetic categorization of art across text, audio, image, and video without explicit labels. Using a self‑supervised framework, the authors embed these modalities into a shared 256‑dimensional space and iteratively cluster the data to uncover aesthetic structure. They compare the AI’s cluster assignments with human affective labels, highlighting divergences and discussing implications for cross‑modal similarity, media organization, and automated labeling.

arXiv Computer Vision
Aug 31

Abstract4D: A Large-Scale Dataset and Framework for Understanding the Visual Language of Abstract Art

Abstract4D is the largest dataset of abstract paintings, containing over 120,000 images with rich metadata and multi‑dimensional prompts that capture perceptual attributes such as form, color, texture, and composition. The dataset is annotated via a hybrid human–VLM pipeline to ensure quality and consistency. Using Abstract4D, the authors analyze the semantic structure of abstract art through large‑scale embedding visualization and establish benchmark tasks for classification, cross‑modal retrieval, and text‑to‑image generation to evaluate AI models’ perception and reproduction of abstract visual language.

By Haowei Zhang, Yuanpei Zhao, Ji-Zhe Zhou, Mao Li
arXiv AI
Sep 10

CS-CLIP: Compositional Scene Graph-guided CLIP for Robust Compositional Reasoning

CS-CLIP is a vision‑language model that improves compositional reasoning by using scene graphs to identify compositional elements and create structured negative examples through selective masking. The approach retains only the most contradictory negatives, encouraging the model to depend on compositional structure instead of surface cues. CS-CLIP achieves state‑of‑the‑art performance on compositional reasoning benchmarks while maintaining strong cross‑modal retrieval and downstream visual reasoning capabilities with fewer training samples.

By SeongJun Jeong, Minjoon Jung, Woo Suk Choi, Youwon Jang, Byoung-Tak Zhang
arXiv Computer Vision
Sep 24

CRISP: Compositional Relations as Invariant Structural Priors for Domain Generalization

CRISP (Compositional Relational Invariance from Spatial Primitives) is an image‑classification framework that decomposes visual recognition into primitive elements and their relational composition. It represents these compositions with soft unary, binary, and ternary predicates over primitive locations and appearance, enabling differentiable spatial and visual alignment learned end‑to‑end. Evaluated on five DomainBed datasets covering style, provenance, and camera‑trap shifts, CRISP achieves new state‑of‑the‑art performance on both benchmarks.

By Dat Nguyen, Duc-Duy Nguyen
arXiv Computer Vision
Aug 27

On the Separation of Human and AI-Generated Images in CLIP Embedding Space

The paper reports a new phenomenon in CLIP embeddings where human and AI‑generated paintings naturally separate along dominant principal directions without any supervised training. The authors investigate this separation by linking embedding directions back to image features using interpretable representations and gradient‑based inversion, finding that the separation is driven by distributed multiscale image structure rather than simple global or local statistics. They also show that small, imperceptible image perturbations can cause large displacements along these directions, highlighting a mismatch between CLIP’s visual evidence and human perception.

By Andrea Asperti