When Do Diffusion Models learn to Generate Multiple Objects?
arXiv:2605. 00273v2 Announce Type: replace-cross Abstract: Text-to-image diffusion models achieve impressive visual fidelity, yet they remain unreliable in multi-object generation.
arXiv:2512. 20666v2 Announce Type: replace-cross Abstract: Text-to-image diffusion models have attracted significant attention for their ability to generate diverse, high-fidelity images.
arXiv:2605. 00273v2 Announce Type: replace-cross Abstract: Text-to-image diffusion models achieve impressive visual fidelity, yet they remain unreliable in multi-object generation.
arXiv:2606. 31699v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) have recently been proposed as interpretable tools for concept-level manipulation, under the assumption that isolated features can serve as controllable intervention points.
arXiv:2603. 00133v2 Announce Type: replace-cross Abstract: Generative models have been shown to "memorize" certain training data, leading to verbatim or near-verbatim generating images, which may cause privacy concerns or copyright infringement.
arXiv:2608. 14172v1 Announce Type: cross Abstract: Text-to-image diffusion models have two major drawbacks that severely limit their practical utility: (1) standard models lack an intrinsic mechanism for continuous, concept-specific guidance (e.
Although text-to-image diffusion models exhibit remarkable generative power, concept erasure techniques are essential for preventing harmful content. Existing adversarial probes evaluate these methods by testing whether erased concepts can still be recovered.
arXiv:2603. 28762v2 Announce Type: replace-cross Abstract: Modern Text-to-Image (T2I) diffusion models have achieved remarkable semantic alignment, yet they often suffer from a significant lack of variety, converging on a narrow set of visual solutions for any given prompt.
arXiv:2608. 05000v1 Announce Type: cross Abstract: Vision offers a critical axis for advancing foundation models, driving a shift towards natively unified multimodal pretraining.
arXiv:2605. 13974v2 Announce Type: replace-cross Abstract: Diffusion Transformers (DiTs) and related flow-based architectures are now among the strongest text-to-image generators, yet the internal mechanisms through which prompts shape image semantics remain poorly understood.
arXiv:2512. 20963v3 Announce Type: replace Abstract: Diffusion models excel at generating high-quality, diverse samples, yet they risk memorizing training data when overfit to the training objective.
arXiv:2606. 25548v1 Announce Type: cross Abstract: Image generative models are trained on massive, largely uncurated internet-scale datasets that contain undesirable visual concepts.
arXiv:2608. 03135v1 Announce Type: cross Abstract: Text-to-image diffusion models can generate individual concepts well, but they often omit or merge concepts incorrectly with multiple concepts.
In human visual perception, uppercase lettering serves as a natural salience cue that captures attention within lowercase text. In this paper, we present a systematic empirical characterization study revealing that Large Language Models (LLMs) exhibit an analogous property: letter casing modulates internal attention allocation.