Hugging Face Trending Papers

Can Text-to-Image Models Draw from the Right Frame of Reference?

Read the original on Hugging Face Trending Papers →

Spatial instruction following has become a crucial requirement for text-to-image (T2I) generation. A common challenge arises when directional expressions are interpreted under different frames of reference.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv Computer Vision
Sep 3

T2LSC-Bench: Benchmarking Localized Semantic Control in Text-to-Image Generation

T2LSC-Bench is a new benchmark for evaluating localized semantic control in text-to-image generation, consisting of 50 seed subjects and 1,200 prompt cases per model, producing 7,160 images across six models. The benchmark measures Text-at-Anchor Accuracy, Semantic Subject Preservation, Semantic Leakage Rate, and Conditional Semantic Leakage Rate using a dual‑branch protocol that combines OCR‑VLM verification with structured VLM semantic judgments. Results show that while accurate text rendering remains high, semantic leakage can increase dramatically under stress‑test conditions, and anti‑leakage prompting can reduce leakage without harming rendering accuracy.

By Yan Wang, Xinyi Hou, Weiguo Lin, Junjun Si, Siwei Ma
arXiv AI
Jul 1

Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking

arXiv:2509. 12046v2 Announce Type: replace-cross Abstract: Although autoregressive (AR) models have demonstrated remarkable success in image generation, extending these models to layout-conditioned generation remains challenging due to the sparse nature of layout conditions and the risk of feature entanglement.

By Zirui Zheng, Takashi Isobe, Tong Shen, Xu Jia, Jianbin Zhao, Xiaomin Li, Mengmeng Ge, Baolu Li, Qinghe Wang, Dong Li, Dong Zhou, Yunzhi Zhuge, Huchuan Lu, Emad Barsoum