arXiv:2605. 21182v2 Announce Type: replace-cross Abstract: Manga is a culturally distinctive multimodal medium and one of the most influential forms of Japanese popular culture.
By Jeonghun Baek, Atsuyuki Miyai, Shota Onohara, Hikaru Ikuta, Kiyoharu Aizawa
LoGAN is a VLM-based agentic framework designed for few-shot multilingual font localization. It takes a handful of glyphs or logo letters and generates complete character sets across many languages, including CJK, by combining a glyph-level diffusion model, style finetuning, spacing/kerning transfer, and texture expansion. The method outperforms specialized font generators and state‑of‑the‑art image editors in glyph fidelity, style, texture, and kerning consistency on datasets covering more than 27 languages.
By Zhuoning Yuan, Ta-Ying Cheng, Benjamin Klein
The paper introduces Synth-JDoc, a synthetic Japanese document image dataset created by rendering text with HTML and CSS to produce multi‑column layouts that include both vertical and horizontal writing styles. Images generated by text‑to‑image models are embedded to enhance visual realism, and noise and degradation filters are applied to improve robustness. Experiments show that fine‑tuning Large Vision Language Models on Synth‑JDoc yields superior performance on reading vertically written Japanese text compared to prior synthetic datasets.
By Keito Sasagawa, Shuhei Kurita, Daisuke Kawahara
arXiv:2605.28173v2 Announce Type: replace
Abstract: End-to-end manga generation is a structured visual storytelling task that requires story decomposition, recurring character and scene grounding, pa...
By Muyao Wang, Zeke Xie, Yanhao Chen, Lixin Xiu, Hideki Nakayama
GlyphAnchor is a new method that improves visual text rendering in image generation and editing models by adding lightweight glyph patch conditions anchored to the target image’s positional encoding. The approach is trained with staged supervised finetuning and text-aware post‑training, and it works with both text‑to‑image and image‑editing diffusion transformers. Experiments on various backbones and the newly introduced InfoTextBench benchmark show that GlyphAnchor consistently enhances text fidelity while maintaining overall image quality, especially for long, complex, or densely arranged text and rare characters.
By Qiang Xiang, Shuang Sun, Binglei Li, Yibo Chen, Xu Tang, Yao Hu, Junping Zhang
arXiv:2609.37141v1 Announce Type: new
Abstract: Semantic typography is a design technique where the visual representation of a word conveys its semantic meaning, while maintaining its legibility. Exi...
By Xinye Yang, Xinding Zhu, Kai Fang, Xinyi Ren, Mengjian Li, Bin Cao, Jiazhou Chen