arXiv Computer Vision

OnomatoBridge: Onomatopoeia Translation and Rendering Pipeline in Manga

OnomatoBridge is a filtering pipeline designed to translate and render onomatopoeia in manga from Japanese to English while preserving the original visual style. The method replaces Japanese onomatopoeia with stylized English text, addressing artifacts and style inconsistencies seen in prior approaches. Experiments on the Manga109 dataset show that OnomatoBridge improves English text correctness by 10–25 points and reduces residual Japanese text by 20–50% compared to baseline image editing models.

arXiv AI
Sep 10

LoGAN: Multilingual Font Localization with Generative Agents

LoGAN is a VLM-based agentic framework designed for few-shot multilingual font localization. It takes a handful of glyphs or logo letters and generates complete character sets across many languages, including CJK, by combining a glyph-level diffusion model, style finetuning, spacing/kerning transfer, and texture expansion. The method outperforms specialized font generators and state‑of‑the‑art image editors in glyph fidelity, style, texture, and kerning consistency on datasets covering more than 27 languages.

By Zhuoning Yuan, Ta-Ying Cheng, Benjamin Klein
arXiv Computation and Language
Aug 31

Synth-JDoc: Synthesizing a Japanese Document Image Dataset for OCR with Diverse Layouts and Embedded Images

The paper introduces Synth-JDoc, a synthetic Japanese document image dataset created by rendering text with HTML and CSS to produce multi‑column layouts that include both vertical and horizontal writing styles. Images generated by text‑to‑image models are embedded to enhance visual realism, and noise and degradation filters are applied to improve robustness. Experiments show that fine‑tuning Large Vision Language Models on Synth‑JDoc yields superior performance on reading vertically written Japanese text compared to prior synthetic datasets.

By Keito Sasagawa, Shuhei Kurita, Daisuke Kawahara
arXiv Computer Vision
Sep 3

GlyphAnchor: Enhancing Visual Text Rendering via Position-Anchored Glyph Priors

GlyphAnchor is a new method that improves visual text rendering in image generation and editing models by adding lightweight glyph patch conditions anchored to the target image’s positional encoding. The approach is trained with staged supervised finetuning and text-aware post‑training, and it works with both text‑to‑image and image‑editing diffusion transformers. Experiments on various backbones and the newly introduced InfoTextBench benchmark show that GlyphAnchor consistently enhances text fidelity while maintaining overall image quality, especially for long, complex, or densely arranged text and rare characters.

By Qiang Xiang, Shuang Sun, Binglei Li, Yibo Chen, Xu Tang, Yao Hu, Junping Zhang
arXiv AI
Jun 12

VDE Bench: Evaluating The Capability of Image Editing Models to Modify Visual Documents

arXiv:2602. 00122v3 Announce Type: replace-cross Abstract: In recent years, image editing models have made significant progress, enabling users to manipulate visual content in a flexible and interactive manner through natural language instructions.

By Hongzhu Yi, Yujia Yang, Yuanxiang Wang, Tong Li, Zhenyu Guan, Tianyu Zong, Jiahuan Chen, Chenxi Bao, Tiankun Yang, Haopeng Jin, Yixuan Yuan, Xinming Wang, Tao Yu, Ruilin Gao, Ruiwen Tao, Haijin Liang, Jin Ma, Jinwen Luo, Yeshani, Xinyu Zuo, Jungang Xu