arXiv Machine Learning

Sequential Multimodal Evidence Optimization for Product Media Ranking in E-Commerce

arXiv:2608. 15662v1 Announce Type: new Abstract: On modern e-commerce stores, customers consume ordered slates of heterogeneous product media, such as images, videos, and 3D renders, before making purchase decisions.

arXiv Machine Learning
Aug 27

DCEO: Direct Causal Effect Optimization for Long-Term User Value Modeling in E-commerce Search

The paper introduces DCEO, a data‑driven framework that learns item‑level proxy scores directly aligned with long‑term user objectives in e‑commerce search. It aggregates these scores into a user‑level metric, measures alignment via relative causal effect, and uses an actor‑critic model to generate context‑dependent fusion weights for multiple objectives. Offline experiments and a 41‑day online A/B test show DCEO improves GMV by 0.36% over traditional proxies.

By Junzhao Zhang, Tao Zhang, Liren Yu, Feiyi Dong, Zhixuan Zhang, Dan Ou, Haihong Tang
arXiv Computer Vision
Sep 18

Grounded Product Understanding in Livestream Videos

The paper introduces GPUB, a large-scale benchmark for grounded product understanding in e‑commerce livestream videos, featuring 3,000 livestreams, 31K fashion products, and multi‑moment temporal annotations. It defines three evaluation tasks, with the main task (GPrU) requiring simultaneous product identification and moment localization. Existing multimodal models perform poorly on GPrU, prompting the authors to develop UniPro, which improves performance by learning product‑aligned, temporally structured representations.

By Xinyu Zhang, Junjie Chen, Jiawei Ge, Qianlong Li, Libin Ma, Baokun Pan, Yahui Luo
arXiv AI
Sep 25

DeGRe: Dense-supervised Generative Reranking for Recommendation

DeGRe is a dense‑supervised generative reranking framework designed to improve multi‑stage recommender systems by addressing label bias and credit assignment issues. It uses an offline Lookahead Evaluator with beam search to generate dense supervision signals, which are distilled into a lightweight Online Generator that can perform efficient greedy decoding at inference time. Experiments show that DeGRe outperforms baselines on public benchmarks and industrial datasets, and it has been successfully deployed on Taobao Flash Shopping to enhance online recommendations.

By Chaotian Song, Jingyao Zhang, Chenghao Chen, Zisen Sang, Dehai Zhao, Guodong Cao, Boxi Wu, Deng Cai, Jia Jia
arXiv AI
Aug 24

TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming

TLive-Omni is an omni‑modal understanding model designed for e‑commerce live streaming, integrating image, video, audio, and text inputs into a unified representation. It introduces Per‑vGrid for timestamped token organization, a three‑stage supervised training pipeline, and a Faithful‑RFT reinforcement fine‑tuning stage to enhance answer faithfulness and expression quality. The model is supported by a scenario‑oriented capability taxonomy and a compact data production engine that generates training signals for tasks such as speech recognition, product visual grounding, and omni‑modal QA, achieving strong performance on live‑commerce benchmarks and good generalization to general tasks.

By Yibo Hu, Yu Qian, Mao Gu, Yingfan Tao, Yuhao Chen, Yongdong Luo, Zhuoqun Liu, Meiguang Jin, Junfeng Ma