arXiv:2609.36598v1 Announce Type: new
Abstract: A video can exhibit convincing motion and photorealism yet fail immediately when visual text collapses. Unlike generic scene content, visual text is un...
By Ziying Zhang, Litao Li, Junchao Liao, Tianyi Zeng, Siyu Zhu, Long Qin, Zhenghao Zhang
GlyphAnchor is a new method that improves visual text rendering in image generation and editing models by adding lightweight glyph patch conditions anchored to the target image’s positional encoding. The approach is trained with staged supervised finetuning and text-aware post‑training, and it works with both text‑to‑image and image‑editing diffusion transformers. Experiments on various backbones and the newly introduced InfoTextBench benchmark show that GlyphAnchor consistently enhances text fidelity while maintaining overall image quality, especially for long, complex, or densely arranged text and rare characters.
By Qiang Xiang, Shuang Sun, Binglei Li, Yibo Chen, Xu Tang, Yao Hu, Junping Zhang
arXiv:2609.40356v1 Announce Type: cross
Abstract: Recent video generation is increasingly realistic and controllable, yet video editing remains less developed, particularly for precise local edits th...
By Xinghao Chen, Xiangbo Gao, Jiongze Yu, Yuheng Wu, Zhengzhong Tu
VTR-Bench is a new benchmark designed to evaluate how well video generation models render text within scenes. It includes 300 prompts across five real-world scenarios such as advertisements and scientific videos, and uses an automated pipeline with human alignment to assess text fidelity and scene/motion requirements. Experiments on 11 state‑of‑the‑art models show that even the best performer has a word error rate of 0.250, underscoring widespread challenges in visual text rendering.
By Yu Huang, Jungang Li, Zhiyuan Wang, Yonghua Hei, Song Dai, Jiayu Yang, Deyuan Liu, Xiang Zheng, Xiaoshuang Shi, Hao Cheng, Kaidi Xu
arXiv:2608. 19637v1 Announce Type: new Abstract: Text editing in product posters entails inserting new text or replacing existing text while preserving product appearance, background content, and global composition.
By Honglie Wang, Jia Sun, Zijun Li, Junlong Wu, Pengcheng Wei, Jiyuan Wang, Yongrui Heng, Boheng Zhang, Huaiqing Wang, Dewen Fan, Qianqian Gan, Fan Yang, Tingting Gao, Yan-Ming Zhang
Edit‑VAR is a training‑free, inversion‑free framework that uses a pretrained visual autoregressive video model for text‑guided video editing. It encodes the source video into multi‑scale discrete tokens and applies probability‑guided conditional token replacement, attention‑guided token‑wise and scale‑aware modulation, and scale‑decoupled generation to preserve source appearance while enabling precise edits. The method also includes residual‑guided token pruning to reduce inference cost, and experimental results show it outperforms existing training‑free video editing methods in fidelity, source preservation, temporal coherence, and efficiency.
By Chongbo Zhao, Jiangming Wang, Xilai Wang, Xinyu Wang, Jingyi Tang, Chunjie Hao, Pengjie Song, Yue Ma
Text editing in product posters entails inserting new text or replacing existing text while preserving product appearance, background content, and global composition. Despite recent progress in instruction-based image editing, general-purpose models remain unreliable in this setting: they often omit or incorrectly render the target text, place it over salient products or pre-existing content, and produce structurally distorted or visually inconsistent glyphs.
arXiv:2607. 19895v1 Announce Type: cross Abstract: Text-guided video editing with diffusion models is impractically slow, hindered by costly multi-step sampling and inversion.
By Habin Lim, Gyeong-Moon Park
ContextAnyone is a context‑aware diffusion framework that treats a reference image as an explicitly preserved appearance anchor rather than a simple conditioning signal. By jointly reconstructing the reference image and generating the target video within a shared diffusion transformer, it provides direct supervision for maintaining identity and fine‑grained appearance throughout denoising. The method introduces asymmetric information flow and Gap‑RoPE positional representations to keep the reference stable while allowing selective access by video tokens, and demonstrates improved identity and appearance consistency on an OpenVid‑HD benchmark.
By Ziyang Mai, Yu-Wing Tai
TextAlign is a post‑training preference‑alignment framework that improves text rendering in large text‑to‑image generative models without changing the generator architecture. It uses a hierarchical vision‑language model to reward global, word, and glyph‑level accuracy, converting binary defect judgments into a scalar preference signal that can be optimized with GRPO or DPO. Experiments on FLUX.1‑dev and Z‑Image‑Turbo demonstrate higher OCR‑based text accuracy while preserving overall generation quality, outperforming several foundation and text‑rendering baselines.
By Mingxuan Cui, Jingpu Yang, Fengxian Ji, Qian Jiang, Zhecheng Shi, Jiaming Wang, Zirui Song, Zhuohan Xie, Fajri Koto, Xiuying Chen
WanPE is a 397‑B parameter prompt‑enhancement model that learns director‑level cinematic planning from 1.05 M real‑world videos. It generates shot‑level cinematic plans through video‑grounded reverse construction and uses Semantic‑Consistency GRPO (SC‑GRPO) to maintain user intent across shots and time. In evaluations, WanPE improves human preference over raw prompts by up to 50.86 points for 30‑second videos and outperforms commercial offerings for shorter durations.
By Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong
EditaLive! is a new real‑time framework for character video editing in live streaming, built on a pretrained image animation model (Wan‑Animate) that separates appearance from motion. It uses the CharEdit‑50K dataset for reference‑frame editing and video reconstruction, and adapts the model from offline bidirectional to causal streaming generation. A self‑rollout distillation strategy compresses the model into a two‑step sampler, employing fixed RoPE, alignment forcing, and first‑frame preserved sparse attention to reduce appearance drift and achieve low‑latency inference while preserving facial expressions.
By Zhiyuan Li, Chi-Man Pun, Peng-Tao Jiang, Bo Li, Xiaodong Cun