arXiv:2512. 08854v3 Announce Type: replace-cross Abstract: It has been hypothesized that achieving the data efficiency of human visual perception requires a generative approach in which internal representations result from inverting a decoder.
By Jack Brady, Bernhard Sch\"olkopf, Thomas Kipf, Simon Buchholz, Wieland Brendel
arXiv:2605. 00273v2 Announce Type: replace-cross Abstract: Text-to-image diffusion models achieve impressive visual fidelity, yet they remain unreliable in multi-object generation.
By Yujin Jeong, Arnas Uselis, Iro Laina, Seong Joon Oh, Anna Rohrbach
arXiv:2606. 13288v1 Announce Type: cross Abstract: Contrastively trained vision-language models like CLIP, have made remarkable progress in learning joint image-text representations, but still face challenges in compositional understanding.
By Wei Li, Zhen Huang, Xinmei Tian
arXiv:2602. 24264v2 Announce Type: replace-cross Abstract: Compositional generalization, the ability to recognize familiar parts in novel contexts, is a defining property of intelligent systems.
By Arnas Uselis, Andrea Dittadi, Seong Joon Oh
arXiv:2607. 02291v1 Announce Type: new Abstract: Conventional reinforcement learning strategies for visual generation typically employ sample-wise reward functions, yet this practice frequently results in reward hacking that degrades image diversity and introduces visual anomalies.
By Ruihang Li, Mengde Xu, Shuyang Gu, Leigang Qu, Fuli Feng, Han Hu, Wenjie Wang
arXiv:2602. 03282v2 Announce Type: replace-cross Abstract: A common assumption in representation learning is that globally well-distributed embeddings support robust and generalizable representations.
By Jiwan Chung, Seon Joo Kim
The paper investigates how long, richly detailed prompts cause modern text-to-image models to lose diversity, even when many visual aspects are unspecified. It introduces PromptMoG, a training‑free method that samples prompt embeddings from a Mixture‑of‑Gaussians distribution to restore diversity while preserving semantic fidelity. The authors also present LPD‑Bench, a benchmark for evaluating fidelity and diversity under long, semantically dense prompts, and demonstrate PromptMoG’s effectiveness on four large diffusion models.
By Bo-Kai Ruan, Teng-Fang Hsiao, Ling Lo, Yi-Lun Wu, Hong-Han Shuai
GeoComposer is a new framework that improves photographic composition by generating textual guidance and visual exemplars for a given image. It uses a geometry-aware representation learning mechanism that incorporates geometric priors from a visual geometry foundation model to maintain global structure and fine-grained correspondences. A reinforcement learning strategy with a hybrid reward optimizes instruction following, aesthetic quality, and geometric consistency, leading to superior results compared to state‑of‑the‑art methods.
By Shuangzhi Li, Qiaoqiao Jia, Xingxin Chen, Guile Wu, Dongfeng Bai
arXiv:2609.35763v3 Announce Type: replace
Abstract: Distributional training provides collective supervision for one-step visual generation by matching real and generated features in frozen representa...
By Chi Zhang, Shi Haoyang, Yueyi Liu, Ruichuan An, Junkang Zhou, Chang Li, Xiuyuan Lu, Yichi Zhang, Bo Wang, Yuhang Wu, Sen Cui, Miao Liu
arXiv:2505.09716v4 Announce Type: replace
Abstract: Out-of-distribution (OOD) generalisation through composition requires a system to discover invariant properties from input-output associations and...
By George Dimitriadis, Spyridon Samothrakis
The paper introduces InterIL, a joint generative model that simultaneously produces a background image and a layout of foreground elements for graphic design templates, addressing the limitations of sequential generation approaches. InterIL connects pretrained image and layout diffusion backbones via a learnable communication module, freezing the backbones to preserve prior knowledge while training only the interaction module. The model also offers a test‑time guidance strategy, enabling users to impose preferences without retraining, and demonstrates superior image, layout, and harmonization quality compared to previous methods.
By Shirong Yang, Bo Yang, Ying Cao
The paper presents a necessary and sufficient condition for provable compositional generalization in neural networks, identifying two key principles: structural alignment and unambiguous minimized representations. It rigorously proves this condition, verifies it in Lean 4, and demonstrates its applicability in few-shot settings, including the SCAN jump task. The authors also outline an algorithmic approach and illustrate it with a minimal example, all derived purely from mathematical analysis without empirical validation.
By Yuanpeng Li