arXiv:2601. 09566v4 Announce Type: replace-cross Abstract: In this work, we study whether rendering Chinese characters as visual glyph images, rather than discrete token IDs as mainstream LLMs do, providing an inductive bias for character-level language modeling.
By Shuyang Xiang, Hao Guan
arXiv:2609.37141v1 Announce Type: new
Abstract: Semantic typography is a design technique where the visual representation of a word conveys its semantic meaning, while maintaining its legibility. Exi...
By Xinye Yang, Xinding Zhu, Kai Fang, Xinyi Ren, Mengjian Li, Bin Cao, Jiazhou Chen
arXiv:2606. 05261v1 Announce Type: cross Abstract: Variable fonts enable continuous variation of glyph geometry along semantic design axes such as weight, width, slant, and optical size.
By Nadav Benedek, Ariel Shamir, Ohad Fried
Few-shot font generation simultaneously requires global structural completeness and fine-grained local style fidelity. Existing methods usually either rely on global content-style modeling, which is robust but imperfectly disentangled, or emphasize component/local modeling, which captures fine details but relies heavily on local priors and reference coverage.
GlyphAnchor is a new method that improves visual text rendering in image generation and editing models by adding lightweight glyph patch conditions anchored to the target image’s positional encoding. The approach is trained with staged supervised finetuning and text-aware post‑training, and it works with both text‑to‑image and image‑editing diffusion transformers. Experiments on various backbones and the newly introduced InfoTextBench benchmark show that GlyphAnchor consistently enhances text fidelity while maintaining overall image quality, especially for long, complex, or densely arranged text and rare characters.
By Qiang Xiang, Shuang Sun, Binglei Li, Yibo Chen, Xu Tang, Yao Hu, Junping Zhang
arXiv:2606. 13382v1 Announce Type: cross Abstract: Few-shot font generation simultaneously requires global structural completeness and fine-grained local style fidelity.
By Zian Yang, Zixin Wang
arXiv:2609.01147v1 Announce Type: cross
Abstract: Text-rich visual inputs require models that can read, retrieve, and compress language directly in pixel space, yet existing pixel-text encoders strug...
By Chaohao Yuan, Ruifeng Yuan, Zhuoxu Huang, Yu Rong, Hong Cheng, Hou Pong Chan, Chenghao Xiao
arXiv:2609.37569v1 Announce Type: new
Abstract: Rendering accurate Chinese text remains challenging for text-to-image models. Existing OCR-based reinforcement-learning rewards compare decoded transcr...
By Yazhen Xie, Xingsong Ye, Zhineng Chen
Rendering accurate Chinese text remains challenging for text-to-image models. Existing OCR-based reinforcement-learning rewards compare decoded transcripts with target strings. Such rewards overlook t...
TextAlign is a post‑training preference‑alignment framework that improves text rendering in large text‑to‑image generative models without changing the generator architecture. It uses a hierarchical vision‑language model to reward global, word, and glyph‑level accuracy, converting binary defect judgments into a scalar preference signal that can be optimized with GRPO or DPO. Experiments on FLUX.1‑dev and Z‑Image‑Turbo demonstrate higher OCR‑based text accuracy while preserving overall generation quality, outperforming several foundation and text‑rendering baselines.
By Mingxuan Cui, Jingpu Yang, Fengxian Ji, Qian Jiang, Zhecheng Shi, Jiaming Wang, Zirui Song, Zhuohan Xie, Fajri Koto, Xiuying Chen
arXiv:2606. 14750v1 Announce Type: cross Abstract: Recent advances in pixel-based text modeling show that representing text as images enables models to exploit visual cues for language understanding.
By Adarsh Arigala, Arjun Gangwar, S Umesh, Yova Kementchedjhieva
OnomatoBridge is a filtering pipeline designed to translate and render onomatopoeia in manga from Japanese to English while preserving the original visual style. The method replaces Japanese onomatopoeia with stylized English text, addressing artifacts and style inconsistencies seen in prior approaches. Experiments on the Manga109 dataset show that OnomatoBridge improves English text correctness by 10–25 points and reduces residual Japanese text by 20–50% compared to baseline image editing models.
By Takara Taniguchi, Wataru Shimoda, Kota Yamaguchi, Hideki Nakayama