CommerceVibe is a system that generates e‑commerce creatives by synthesizing executable HTML/CSS code conditioned on product images, design requirements, and product information. It uses dual‑feedback reinforcement learning, combining rule‑based checks for text readability, product visibility, and layout validity with visual feedback from a vision‑language model that evaluates perceptual and commercial aspects. After fine‑tuning a large language model on 28,000 examples and applying dual‑feedback reinforcement learning, CommerceVibe achieves a weighted score of 94.0/100 on a 1,300‑case benchmark, outperforming both its SFT‑only counterpart and external models, and is validated by expert blind evaluations.
By Yajiao Xu, Jin Zhang, Jiangbo Ai, Tao Jiang, Mo Xu, Lina Huang, Chengfu Huo
RILA is an execution‑driven agent that integrates browser rendering into the generation loop for interactive web development. It uses an Action Interaction Verification module to replay reference interactions on generated pages, collecting execution‑aware observations, and an Execution‑aware Rendering Score to jointly assess interaction correctness and visual fidelity during iterative optimization. A data synthesis pipeline further augments training data, enabling RILA to significantly improve interaction and visual quality across foundation models, even outperforming larger one‑shot generators.
By Yilong Guo, Hanqi Chen, Zixiao Ye, Guanzhong Wang, Chen Yu, Zeyu Chen
arXiv:2606. 00154v1 Announce Type: cross Abstract: Recent advancements in multimodal large language models (MLLMs) have achieved remarkable progress in multimodal reasoning and code generation, catalyzing a new paradigm for front-end development.
By Fan Wu, Lishuai Dong, Cuiyun Gao, Yujia Chen, Yiming Huang, Yang Xiao, Qing Liao
arXiv:2606. 19103v1 Announce Type: cross Abstract: Recent advances in instruction-based image editing have enabled models to perform complex visual edits from natural language instructions.
By Mukund Khanna, Raj Singh Yadav, Kunal Singh
Cross-border e-commerce image translation is essential for global retail, where product images, banners, and detail pages need to be produced in different languages. Existing methods struggle to achieve accurate translation, faithful visual identity preservation, and easy-to-edit outputs, simultaneously.
arXiv:2608. 03691v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) are increasingly used to translate webpage screenshots into front-end code, but repeated UI patterns may sway them toward visually incorrect yet pattern-consistent outputs.
By Khai-Nguyen Nguyen, Oscar Chaparro, Antonio Mastropaolo
arXiv:2608. 11907v2 Announce Type: replace-cross Abstract: As Large Vision-Language Models increasingly aim to integrate visual generation and understanding within a single parameter space, evaluating such structural unification in a cohesive manner remains a critical challenge.
By Hao Zhang, Jiaxin Qi, Zhijiang Tang, Jianqiang Huang
As Large Vision-Language Models increasingly aim to integrate visual generation and understanding within a single parameter space, evaluating such structural unification in a cohesive manner remains a critical challenge. Current evaluation protocols predominantly treat generative and discriminative capabilities as separate tasks, leaving a gap in system-level evaluation for unified multimodal models (UMMs).
arXiv:2608.23302v1 Announce Type: new
Abstract: Fashion complementary image generation (CIG) aims to create garments that stylistically match a seed item based on user intent, making it a natural mul...
By Matteo Attimonelli, Claudio Pomo, Alessandro De Bellis, Danilo Danese, Dietmar Jannach, Tommaso Di Noia
arXiv:2607. 10079v1 Announce Type: new Abstract: Digital Adoption Platforms (DAPs) are embedded overlays widely used on web systems to guide users through operations inside a page, helping them get started with unfamiliar interfaces quickly.
By Chengguang Gan, Hanjun Wei, Yunhao Liang, Zhixi Cai, Qinghao Zhang, Shiwen Ni
arXiv:2608. 11907v1 Announce Type: cross Abstract: As Large Vision-Language Models increasingly aim to integrate visual generation and understanding within a single parameter space, evaluating such structural unification in a cohesive manner remains a critical challenge.
By Hao Zhang, Jiaxin Qi, Zhijiang Tang, Jianqiang Huang
Fashion complementary image generation (CIG) aims to create garments that stylistically match a seed item based on user intent, making it a natural multimodal grounding problem where models must inter...