arXiv:2608. 19637v1 Announce Type: new Abstract: Text editing in product posters entails inserting new text or replacing existing text while preserving product appearance, background content, and global composition.
By Honglie Wang, Jia Sun, Zijun Li, Junlong Wu, Pengcheng Wei, Jiyuan Wang, Yongrui Heng, Boheng Zhang, Huaiqing Wang, Dewen Fan, Qianqian Gan, Fan Yang, Tingting Gao, Yan-Ming Zhang
LoGAN is a VLM-based agentic framework designed for few-shot multilingual font localization. It takes a handful of glyphs or logo letters and generates complete character sets across many languages, including CJK, by combining a glyph-level diffusion model, style finetuning, spacing/kerning transfer, and texture expansion. The method outperforms specialized font generators and state‑of‑the‑art image editors in glyph fidelity, style, texture, and kerning consistency on datasets covering more than 27 languages.
By Zhuoning Yuan, Ta-Ying Cheng, Benjamin Klein
Text editing in product posters entails inserting new text or replacing existing text while preserving product appearance, background content, and global composition. Despite recent progress in instruction-based image editing, general-purpose models remain unreliable in this setting: they often omit or incorrectly render the target text, place it over salient products or pre-existing content, and produce structurally distorted or visually inconsistent glyphs.
arXiv:2605. 16409v3 Announce Type: replace-cross Abstract: Optical character recognition (OCR) and multilingual scene-text understanding remain challenging for multimodal large language models (MLLMs), particularly in real-world images containing small or degraded text, cluttered layouts, occlusion, handwriting, and complex typography.
By Qinwu Xu, Yifan Jiang, Haoyu Ren
GlyphAnchor is a new method that improves visual text rendering in image generation and editing models by adding lightweight glyph patch conditions anchored to the target image’s positional encoding. The approach is trained with staged supervised finetuning and text-aware post‑training, and it works with both text‑to‑image and image‑editing diffusion transformers. Experiments on various backbones and the newly introduced InfoTextBench benchmark show that GlyphAnchor consistently enhances text fidelity while maintaining overall image quality, especially for long, complex, or densely arranged text and rare characters.
By Qiang Xiang, Shuang Sun, Binglei Li, Yibo Chen, Xu Tang, Yao Hu, Junping Zhang
arXiv:2609.36598v1 Announce Type: new
Abstract: A video can exhibit convincing motion and photorealism yet fail immediately when visual text collapses. Unlike generic scene content, visual text is un...
By Ziying Zhang, Litao Li, Junchao Liao, Tianyi Zeng, Siyu Zhu, Long Qin, Zhenghao Zhang
MUDIDI is a two-stage framework designed to digitize multilingual dictionaries that are currently only available as scanned images. The first stage assesses character recognition and markup preservation, while the second stage segments dictionary entries and maps them into the SIL Multi-Dictionary Formatter schema. The authors also release a dataset of 30 annotated dictionaries and benchmark OCR, LLM, and VLM systems, finding that LLMs generally outperform others and that providing additional context improves digitization quality.
By David Setiawan, Temuulen Khishigsuren, Milind Agarwal, Pagnarith Pit, Aso Mahmudi, Ekaterina Vylomova
arXiv:2609.37569v1 Announce Type: new
Abstract: Rendering accurate Chinese text remains challenging for text-to-image models. Existing OCR-based reinforcement-learning rewards compare decoded transcr...
By Yazhen Xie, Xingsong Ye, Zhineng Chen
arXiv:2604.02103v3 Announce Type: replace-cross
Abstract: Realistic online handwriting depends not only on individual character shapes, but also on how a writer connects, spaces, and aligns adjacent...
By Jinsu Shin, Sungeun Hong, JinYeong Bak
Rendering accurate Chinese text remains challenging for text-to-image models. Existing OCR-based reinforcement-learning rewards compare decoded transcripts with target strings. Such rewards overlook t...
arXiv:2604. 05853v3 Announce Type: replace Abstract: Modern text-to-image (T2I) models can now render legible, paragraph-length text, enabling a fundamentally new class of misuse.
By Zonghao Ying, Haowen Dai, Lianyu Hu, Zonglei Jing, Quanchen Zou, Yaodong Yang, Aishan Liu, Xianglong Liu
The paper introduces Synth-JDoc, a synthetic Japanese document image dataset created by rendering text with HTML and CSS to produce multi‑column layouts that include both vertical and horizontal writing styles. Images generated by text‑to‑image models are embedded to enhance visual realism, and noise and degradation filters are applied to improve robustness. Experiments show that fine‑tuning Large Vision Language Models on Synth‑JDoc yields superior performance on reading vertically written Japanese text compared to prior synthetic datasets.
By Keito Sasagawa, Shuhei Kurita, Daisuke Kawahara