arXiv AI By Zhihong Liu, Siqi Kou, Zheng Li, Ye Ma, Quan Chen, Peng Jiang, Kai Yu, Zhijie Deng

ProductWebGen: Benchmarking Multimodal Product Webpage Generation

Read the original on arXiv AI →

arXiv:2606. 01022v1 Announce Type: cross Abstract: Crafting a product display webpage from a source product image, along with layout and visual content instructions, holds significant practical value for domains such as marketing, advertising, and E-commerce.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Aug 31

CommerceVibe: Learning to Design E-Commerce Creatives as Executable Visual Code via Dual-Feedback Reinforcement Learning

CommerceVibe is a system that generates e‑commerce creatives by synthesizing executable HTML/CSS code conditioned on product images, design requirements, and product information. It uses dual‑feedback reinforcement learning, combining rule‑based checks for text readability, product visibility, and layout validity with visual feedback from a vision‑language model that evaluates perceptual and commercial aspects. After fine‑tuning a large language model on 28,000 examples and applying dual‑feedback reinforcement learning, CommerceVibe achieves a weighted score of 94.0/100 on a 1,300‑case benchmark, outperforming both its SFT‑only counterpart and external models, and is validated by expert blind evaluations.

By Yajiao Xu, Jin Zhang, Jiangbo Ai, Tao Jiang, Mo Xu, Lina Huang, Chengfu Huo
arXiv Computer Vision
Sep 3

Rendering-in-the-Loop: An Execution-Driven Agent for Interactive Web Development

RILA is an execution‑driven agent that integrates browser rendering into the generation loop for interactive web development. It uses an Action Interaction Verification module to replay reference interactions on generated pages, collecting execution‑aware observations, and an Execution‑aware Rendering Score to jointly assess interaction correctness and visual fidelity during iterative optimization. A data synthesis pipeline further augments training data, enabling RILA to significantly improve interaction and visual quality across foundation models, even outperforming larger one‑shot generators.

By Yilong Guo, Hanqi Chen, Zixiao Ye, Guanzhong Wang, Chen Yu, Zeyu Chen
Hugging Face Trending Papers
Aug 17

TransAnyText: Translating Arbitrary Text in E-commerce Images via Structured Visual Generation

Cross-border e-commerce image translation is essential for global retail, where product images, banners, and detail pages need to be produced in different languages. Existing methods struggle to achieve accurate translation, faithful visual identity preservation, and easy-to-edit outputs, simultaneously.