arXiv AI

Balancing Trial and Reorder: A Hybrid Sequential Transformer-GBDT Ranker for On-Demand Delivery

arXiv Computation and Language
6d ago

ZooWork-ShopRanker: An Open, Preference-Aligned E-Commerce Reranker

ZooWork-ShopRanker is a family of open e‑commerce rerankers (0.6B, 4B, and 8B) that align with human shopping preferences by using large language models as preference oracles to generate training pairs. The flagship 8B model serves as a teacher for the smaller 4B and 0.6B models, which are further refined on judged pairs. A new benchmark, ShopRank‑Bench, contains ~10,000 private‑traffic preference pairs and shows that all ZooWork models outperform the strongest open reranker baseline and their own un‑aligned versions.

By Siqiao Xue, Shuxuan Liu, Ning Hu
arXiv AI
Sep 1

Beyond Ranking Accuracy: Evaluating LLM-Cited Feature Rationales for Next Basket Repurchase Recommendation

arXiv:2608.30333v1 Announce Type: cross Abstract: Next-basket repurchase recommendation is commonly formulated as a ranking task: given a customer's purchase history, the system ranks previously purc...

By Yanan Cao, Anay Dombe, Murali Mohana Krishna Dandu, Shreeranjani Srirangamsridharan, Sinduja Subramaniam, Yogananth Mahalingam, Evren Korpeoglu, Kannan Achan
arXiv AI
Sep 2

From Language to Behavior: Scaling Sequence Transformers for Industrial Recommendation Ranking with Rec-Native Designs

The paper introduces ReST, a recommendation‑native Transformer scaling framework designed to handle noisy, irregular, and sparsely supervised user behavior sequences in production ranking. ReST employs a dual‑gated attention encoder with rotary positional and temporal embeddings, and a lightweight cross decoder that decouples heavy encoding from fast decoding, enabling efficient compute‑once, decode‑many‑times ranking. Experiments on industrial and public benchmarks show that ReST outperforms traditional Transformer blocks, achieving higher accuracy and consistent scaling across sequence length, depth, and width, and a one‑week online A/B test on a production advertising platform yielded a 1.31% AUC lift and an 11.93% increase in a core revenue metric within a 50 ms P99 latency budget.

By Jie Chen, Xiangqian Yu, Yanchao Lian, Tan Lu, Run Yang, Zhengchun Shang, Xing Wang, Cheng Chen, Ke Hu, Qiang Li, Tianjiu Yin, Xiaobing Liu
arXiv Machine Learning
Sep 1

E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation

E-Commerce Bench is an open‑source benchmark that simulates a year‑long e‑commerce operation, requiring LLM agents to manage multiple online stores, negotiate with suppliers, optimize sales, fulfill orders, handle returns, and manage cash flow. The environment uses real product and supplier data, a calendar of promotions and shocks, and deterministic customer and negotiation models to enable reproducible evaluation. The study evaluates 18 state‑of‑the‑art models across seven metrics, finding no single model dominates, with GPT‑5.6 Sol achieving the highest year‑end assets but lagging in fraud avoidance and operational efficiency.

By Wei Fan, Xinjie Shen, Xudong Guo, Jianhong Tu, Yang Su, Yinger Zhang, Lianghao Deng, Fengyu Wang, Baohua Dong, Yangqiu Song, Dayiheng Liu