arXiv Computer Vision By Yan Wang, Yanzu Wang, Maitreyee Joshi, Samiha Sadeka, Partho Hassan, Reza Jarral, Sayeef Abdullah, Sabit Hassan

Moonworks Lunara: Modeling Artistic Intelligence

Read the original on arXiv Computer Vision →

Moonworks Lunara is a text‑to‑image model that defines Artistic Intelligence as exploration‑driven world realization, preserving semantic, artistic, and compositional structure. It uses a Diffusion Mixture Transformer architecture and a training algorithm that iteratively refines the data distribution with informative samples and human‑created art. Benchmarks show Lunara ranks first in aesthetic quality and second in emotional resonance against seven other image‑generation models, while maintaining a sub‑10B parameter size and sub‑10‑second inference latency.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv AI
Sep 17

CompArt: Operationalizing Aesthetic Alignment in Text-to-Image Generation via Principles of Art

CompArt introduces a new approach to aesthetic alignment in text-to-image generation by using the Principles of Art (PoA) such as Balance, Rhythm, and Emphasis to define explicit compositional constraints. The authors create a large dataset of 80,032 WikiArt images, each annotated with PoA analyses generated by a multimodal LLM, and present ArtDapter, a lightweight adapter that steers a pretrained diffusion model along ten PoA dimensions while preserving semantic fidelity. Experiments demonstrate that CompArt outperforms strong baselines in adhering to PoA controls under a dual evaluation protocol.

By Zhe Jin, Tat-Seng Chua
Hugging Face Trending Papers
Aug 3

MIEScore: Human-Aligned Evaluation for Multi-Source Image Editing

Recent advances in unified multimodal models have significantly improved text-guided image editing abilities. In particular, models such as Nano-Banana-Pro and GPT-Image-2 demonstrate emerging capabilities in multi-source image editing (MIE), including tasks such as object synthesis, person-background composition, and cross-image style fusion.

Hugging Face Trending Papers
Aug 14

CRAFT: Constrained Reward via Attention Fine-Tuning for Subject Personalization without Composed Targets

Subject-driven image personalization---generating new images that preserve the identity of one or several reference subjects in novel scenes---is a foundational capability for modern visual content creation. It is currently dominated by generalized methods that fine-tune a pretrained multimodal diffusion transformer (MMDiT) on hundreds of thousands to millions of paired \emph{(reference, composed-target)} examples, where each composed target is a synthesized image of the subject in a novel scene.

arXiv AI
3d ago

ReGain: Restoring Subject Fidelity in Personalization on Synthetic Images

ReGain is a training‑free correction that improves subject fidelity in text‑to‑image diffusion models personalized with synthetic images. The authors show that fine‑tuning on synthetic images degrades fidelity due to inflated classifier‑free guidance, especially at high frequencies. ReGain measures this inflation per frequency band and scales it down during sampling, closing 51‑64% of the fidelity gap on Stable Diffusion v1.5 and improving performance on SDXL and SD 3.5 while preserving text alignment.

By Shubhang Bhatnagar, Ishan Bhatnagar, Viraj Shah, Narendra Ahuja