arXiv AI

When Pretty Isn't Useful: Investigating Why Modern Text-to-Image Models Fail as Reliable Training Data Generators

arXiv:2602. 19946v5 Announce Type: replace-cross Abstract: Recent text-to-image (T2I) diffusion models produce visually stunning images and demonstrate excellent prompt following.

Hugging Face Trending Papers
Jun 23

Benchmarking the Alignment of Data-Quality Metrics, Human Judgment and Land-Cover Segmentation Performance for Earth Observation

Volume and quality of datasets are crucial for deep learning model training, yet they are often constrained by availability and data acquisition costs. Synthetic data augmentation can extend existing datasets with realistic images, and the quality of these images is generally assessed through fidelity metrics such as FID, KID, IS, LPIPS and SSIM that measure structural or distributional similarity.

arXiv Computer Vision
Sep 14

Unified Text-Image Generation with Weakness-Targeted Post-Training

The paper introduces a post‑training approach that enables a single inference process to transition from text reasoning to image synthesis, eliminating the need for explicit modality switching. Using the 14B BAGEL model, the authors demonstrate that targeted post‑training data and reward‑weighted training improve multimodal image generation across four independent T2I benchmarks. The study highlights the benefits of joint text‑image generation and strategic data selection for enhancing T2I performance.

By Jiahui Chen, Philippe Hansen-Estruch, Xiaochuang Han, Yushi Hu, Emily Dinan, Amita Kamath, Michal Drozdzal, Reyhane Askari-Hemmat, Luke Zettlemoyer, Marjan Ghazvininejad
arXiv Machine Learning
Aug 28

A Framework for Low-Effort Training Data Generation for Urban Semantic Segmentation

The paper introduces a framework that adapts a diffusion model to a target urban domain using only imperfect pseudo‑labels, enabling the generation of high‑fidelity, target‑aligned images from semantic maps of any synthetic dataset. By filtering poor generations, correcting image‑label misalignments, and standardising semantics, the method transforms low‑effort synthetic data into competitive real‑domain training sets. Experiments on five synthetic and two real datasets show up to +8.0 %pt mIoU improvement over state‑of‑the‑art translation methods, demonstrating that rapidly constructed synthetic datasets can match the performance of high‑effort, manually designed ones.

By Damjan Kal\v{s}an, Denis Zavadski, Tim K\"uchler, Haebom Lee, Stefan Roth, Carsten Rother
arXiv AI
Jul 13

Video Generation Models are General-Purpose Vision Learners

arXiv:2607. 09024v1 Announce Type: cross Abstract: Driven by next-token prediction, NLP shifted from task-specific models into powerful generalist foundation models.

By Letian Wang, Chuhan Zhang, Rishabh Kabra, Jasper Uijlings, Steven Waslander, Andrew Zisserman, Joao Carreira, Kaiming He, Misha Andriluka, Eduard Gabriel Bazavan, Andrei Zanfir, Cristian Sminchisescu