arXiv AI

What Drives Compositional Generalization in Visual Generative Models? The Importance of Continuous Training Objectives

arXiv Machine Learning
Jul 7

Is Generation Required for Data-Efficient Perception?

arXiv:2512. 08854v3 Announce Type: replace-cross Abstract: It has been hypothesized that achieving the data efficiency of human visual perception requires a generative approach in which internal representations result from inverting a decoder.

By Jack Brady, Bernhard Sch\"olkopf, Thomas Kipf, Simon Buchholz, Wieland Brendel
arXiv Machine Learning
Jul 3

Optimizing Visual Generative Models via Distribution-wise Rewards

arXiv:2607. 02291v1 Announce Type: new Abstract: Conventional reinforcement learning strategies for visual generation typically employ sample-wise reward functions, yet this practice frequently results in reward hacking that degrades image diversity and introduces visual anomalies.

By Ruihang Li, Mengde Xu, Shuyang Gu, Leigang Qu, Fuli Feng, Han Hu, Wenjie Wang
arXiv Computer Vision
Sep 3

Diversifying Long Prompt Image Generation through Structured Prompt Embedding Space Sampling

The paper investigates how long, richly detailed prompts cause modern text-to-image models to lose diversity, even when many visual aspects are unspecified. It introduces PromptMoG, a training‑free method that samples prompt embeddings from a Mixture‑of‑Gaussians distribution to restore diversity while preserving semantic fidelity. The authors also present LPD‑Bench, a benchmark for evaluating fidelity and diversity under long, semantically dense prompts, and demonstrate PromptMoG’s effectiveness on four large diffusion models.

By Bo-Kai Ruan, Teng-Fang Hsiao, Ling Lo, Yi-Lun Wu, Hong-Han Shuai
arXiv Computer Vision
Sep 23

GeoComposer: Geometry-Grounded Photographic Composition Instruction

GeoComposer is a new framework that improves photographic composition by generating textual guidance and visual exemplars for a given image. It uses a geometry-aware representation learning mechanism that incorporates geometric priors from a visual geometry foundation model to maintain global structure and fine-grained correspondences. A reinforcement learning strategy with a hybrid reward optimizes instruction following, aesthetic quality, and geometric consistency, leading to superior results compared to state‑of‑the‑art methods.

By Shuangzhi Li, Qiaoqiao Jia, Xingxin Chen, Guile Wu, Dongfeng Bai
arXiv Computer Vision
Sep 11

Learning Interaction between Image and Layout Priors for Joint Image-Layout Generation in Design Templates

The paper introduces InterIL, a joint generative model that simultaneously produces a background image and a layout of foreground elements for graphic design templates, addressing the limitations of sequential generation approaches. InterIL connects pretrained image and layout diffusion backbones via a learnable communication module, freezing the backbones to preserve prior knowledge while training only the interaction module. The model also offers a test‑time guidance strategy, enabling users to impose preferences without retraining, and demonstrates superior image, layout, and harmonization quality compared to previous methods.

By Shirong Yang, Bo Yang, Ying Cao
arXiv AI
Sep 10

A Theoretical Analysis of Provable Compositional Generalization in Neural Networks: A Necessary and Sufficient Condition

The paper presents a necessary and sufficient condition for provable compositional generalization in neural networks, identifying two key principles: structural alignment and unambiguous minimized representations. It rigorously proves this condition, verifies it in Lean 4, and demonstrates its applicability in few-shot settings, including the SCAN jump task. The authors also outline an algorithmic approach and illustrate it with a minimal example, all derived purely from mathematical analysis without empirical validation.

By Yuanpeng Li