arXiv AI By Mukund Khanna, Raj Singh Yadav, Kunal Singh

ProductConsistency: Improving Product Identity Preservation in Instruction-Based Image Editing via SFT and RL

Read the original on arXiv AI →

arXiv:2606. 19103v1 Announce Type: cross Abstract: Recent advances in instruction-based image editing have enabled models to perform complex visual edits from natural language instructions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Aug 21

TextRefine: Improving Textual Fidelity, Spatial Placement, and Glyph Rendering for Text Editing in Product Posters

arXiv:2608. 19637v1 Announce Type: new Abstract: Text editing in product posters entails inserting new text or replacing existing text while preserving product appearance, background content, and global composition.

By Honglie Wang, Jia Sun, Zijun Li, Junlong Wu, Pengcheng Wei, Jiyuan Wang, Yongrui Heng, Boheng Zhang, Huaiqing Wang, Dewen Fan, Qianqian Gan, Fan Yang, Tingting Gao, Yan-Ming Zhang
Hugging Face Trending Papers
Aug 20

TextRefine: Improving Textual Fidelity, Spatial Placement, and Glyph Rendering for Text Editing in Product Posters

Text editing in product posters entails inserting new text or replacing existing text while preserving product appearance, background content, and global composition. Despite recent progress in instruction-based image editing, general-purpose models remain unreliable in this setting: they often omit or incorrectly render the target text, place it over salient products or pre-existing content, and produce structurally distorted or visually inconsistent glyphs.

arXiv Computer Vision
Aug 27

RefVideo-6M: A Reliable Reference-Based Dataset for Instructional Video Editing

RefVideo-6M is a new large-scale reference-guided editing dataset that includes 5 million video editing samples and 1 million image editing samples, each paired with about 6 million visual references. The dataset is constructed to avoid artifacts by using real, artifact‑free videos as targets and filtering input conditions with multiple editing experts, thereby providing reliable supervision. It enables models to learn fine‑grained visual correspondence beyond text‑only instructions and supports the training of a reference‑guided video editing model, Ref‑MoT, which shows improved visual quality, controllability, and reference consistency.

By Bojia Zi, Xiaoyan Yang, Yu Zhou, Ruijie Sun, Lihan Zhang, Bin Liang, Kam-Fai Wong, Haibin Huang, Chi Zhang, Xuelong Li
Hugging Face Trending Papers
Aug 3

MIEScore: Human-Aligned Evaluation for Multi-Source Image Editing

Recent advances in unified multimodal models have significantly improved text-guided image editing abilities. In particular, models such as Nano-Banana-Pro and GPT-Image-2 demonstrate emerging capabilities in multi-source image editing (MIE), including tasks such as object synthesis, person-background composition, and cross-image style fusion.