arXiv Machine Learning

DODA: A Database of Datasets for Aesthetics Research

arXiv:2608. 00089v1 Announce Type: cross Abstract: With rapid growth in the fields of empirical and computational aesthetics we have seen a vast increase in large image datasets annotated for aesthetics.

arXiv Computer Vision
Sep 1

DisciplineGen-1M: A Large-Scale Dataset for Multidisciplinary Visual Generation and Editing

arXiv:2607.02290v2 Announce Type: replace Abstract: Recent image generation and editing models can produce visually appealing natural images, yet they remain unreliable when the target image is a kno...

By Zhaokai Wang, Mingxin Liu, Zirun Zhu, Ziqian Fan, Yiguo He, Mohan Zhang, Leyao Gu, Yan Li, Xiangyu Zhao, Ning Liao, Shaofeng Zhang, Xuanhe Zhou, Zhihang Zhong, Xue Yang
arXiv Computer Vision
Sep 2

MegaStyle++: Scaling Image Style Space through Hierarchical Style Definition

MegaStyle++ introduces a hierarchical definition of image style, ranging from overall style identity to fine‑grained visual attributes, to provide a more structured, transferable, and interpretable representation. Using this definition, the authors refined the MegaStyle annotation pipeline and released MegaStyle++‑8M, a dataset with 150K style identities, 1M fine‑grained prompts, and 8M stylized images. Analyses show that the hierarchical approach expands style diversity and semantic breadth while accurately capturing the intrinsic visual style of reference images.

By Junyao Gao, Sibo Liu, Jiaxing Li, Yanan Sun, Weidong Zhang, Cairong Zhao, Jun Zhang
Hugging Face Trending Papers
Jun 29

SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation

While Text-to-Image (T2I) models have shown remarkable success in generating photorealistic visual content, they still struggle with the rigorous semantic alignment and logical reasoning required for scientific imagery. Inspired by Peirce's Semiotic Triad, we introduce Scientific Image Reasoning (SciIR), a comprehensive resource for training and evaluation of scientific image generation.

arXiv AI
Aug 11

Beyond Pixels: Benchmarking and Reward-Based Assessing Framework for Visual Spatial Aesthetics

arXiv:2512. 05098v2 Announce Type: replace-cross Abstract: In recent years, Image Quality Assessment (IQA) for AI-generated images (AIGI) has advanced rapidly; however, existing methods primarily target portraits and artistic images, lacking a systematic evaluation of interior scenes.

By Yuan Gao, Jin Song, Yiyun Fei, Gongzhe Li, Ruigao Yang
arXiv AI
2d ago

Personalized Image Generation with Reasoning and Reflection

arXiv:2610.00737v1 Announce Type: cross Abstract: Personalized image generation has remained narrowly focused on conditional synthesis from curated visual exemplars, rather than capturing who a user...

By Bo Ni, Ngoc N. Tran, Qinwen Ge, Franck Dernoncourt, Seunghyun Yoon, Samyadeep Basu, Sungchul Kim, Puneet Mathur, Nedim Lipka, Tong Yu, Yu Wang, Ryan A. Rossi, Tyler Derr
arXiv Machine Learning
Aug 28

How AI Experiences Art: Emergent Aesthetic Structure in a Self-Supervised Multimodal Embedding Space

The paper introduces a self‑supervised framework that maps text, audio, image, and video into a shared 256‑dimensional embedding space and uses iterative clustering to uncover aesthetic structure. It examines how AI’s cluster assignments diverge from human affective labels on a weakly supervised multimodal dataset. The study highlights implications for cross‑modal similarity, media organization for Retrieval‑Augmented Generation, and automated data labeling.

By Corey D. C. Heath
Hugging Face Trending Papers
Aug 27

How AI Experiences Art: Emergent Aesthetic Structure in a Self-Supervised Multimodal Embedding Space

The paper explores how AI can develop its own aesthetic categorization of art across text, audio, image, and video without explicit labels. Using a self‑supervised framework, the authors embed these modalities into a shared 256‑dimensional space and iteratively cluster the data to uncover aesthetic structure. They compare the AI’s cluster assignments with human affective labels, highlighting divergences and discussing implications for cross‑modal similarity, media organization, and automated labeling.