arXiv Computation and Language

Scaling E-Commerce Attribute Extraction with Parallel Decoding

The paper presents a two‑stage large language model pipeline for e‑commerce attribute extraction. First, it discovers a compact, ranked set of purchase‑discriminative attributes for each product category; second, it extracts those attribute values from catalog text using a fine‑tuned Qwen3‑4B model with Hyper‑Parallel Decoding. The approach attains 85% extraction accuracy while cutting inference costs by 92%, enabling scalable production use and producing consistent, comparable product knowledge bases.

arXiv AI
Jun 24

AI-PAVE-Br: Leveraging Large Language Models for Enhanced Product Attribute Value Extraction through a Golden Set Approach

arXiv:2606. 24655v1 Announce Type: cross Abstract: The explosive growth and complexity of product data within the dynamic Brazilian e-commerce landscape demand robust and specialized methods for structured information extraction.

By Murilo Gazzola, Hugo Gobato Souto, Samuel Silva, J\'ulia Schubert Peixoto, Felipe Siqueira, Andr\'e Luis Pedroso de Morais, Caio Gomes
arXiv Machine Learning
Jun 2

Semantic Retrieval for Product Search in E-Commerce

arXiv:2606. 01504v1 Announce Type: cross Abstract: Semantic retrieval in e-commerce must handle short, noisy, and colloquial queries over large product catalogs with fine-grained attribute distinctions.

By Nikhil Kothari, Saksham Samdani, Ritam Mallick, Praveen Gupta, Ankit Vijay, Surender Kumar
arXiv Computation and Language
6d ago

REALMS: An AI-Assistant Conversational System for Real-Time Exact Audience Sizing over High-Dimensional Nested Profiles

REALMS is a conversational system that provides real‑time, exact audience sizing for digital marketers. It uses embedding‑based vector search to retrieve relevant categorical attributes, an LLM‑powered NL2SQL pipeline for accurate query generation over complex nested schemas, and schema standardization for industry‑agnostic deployment. Evaluations on real enterprise data show high recall, accurate SQL execution, and low latency, enabling interactive audience insights that previously required hours.

By Haixu Ma, Aditya Bansal, Shubham Lohiya, Sumit Ranjan
arXiv AI
Aug 24

One Hierarchy, Two Systems: Semantic Product IDs for Discovery-Surface Ranking and Search-Page Query Reformulation

The paper proposes a single hierarchical Semantic ID (SID) system to unify product identification across multiple merchants in e-commerce. By learning SID representations from product content, the authors demonstrate that ranking algorithms can aggregate consumer affinity and product performance over SID prefixes, improving offline relevance and online engagement. For query reformulation, SID concepts guide navigation and refinement, yielding better intent preservation and higher-quality suggestions compared to taxonomy or raw query transitions.

By Steven Xu, Sanjyot Thete, Saathvik Dirisala, Raghav Saboo, Nimesh Sinha, Leo Shao, Elyse Winer, Sudeep Das, Martin Wang, Kyle MacDonald
arXiv AI
Jun 26

AIGP: An LLM-Based Framework for Long-Term Value Alignment in E-Commerce Pricing

arXiv:2606. 26787v1 Announce Type: cross Abstract: Traditional dynamic pricing models in large-scale e-commerce suffer from limited interpretability, poor utilization of unstructured information, and misalignment with long-term business objectives such as cumulative Gross Merchandise Value (GMV), Return on Investment (ROI) and milestone achievement.

By Chennan Ma, Yanning Zhang, Siqi Hong, Xiuchong Wang, Fei Xiao, Keping Yang
arXiv Machine Learning
Jul 14

Serving the Long Tail: Training-Free LLM Candidate Generation for Vacation Rental Marketplaces

arXiv:2607. 09877v1 Announce Type: new Abstract: Vacation rental marketplaces face a structural imbalance on the supply side: a small fraction of properties receive most user interactions, while the long tail of new, niche, and seasonal listings generates too little behavioral signal for collaborative filtering to serve effectively.

By Syed Mohammed Arshad Zaidi, Eric Rincon, Shayan Hassantabar
arXiv Machine Learning
1d ago

RPTune: Learned Context Curation for LLM Catalog Search

RPTune is an end‑to‑end framework that improves in‑context catalog search for small merchant businesses by learning to curate product catalogs and fine‑tuning large language models (LLMs) with catalog‑grounded supervision. It uses an encoder‑reorganizer curator to order and prune products based on LLM feedback, and then applies context‑relative rewards during LLM post‑training. Across seven real merchants and 100 complex conversational queries per merchant, RPTune boosts search accuracy by up to 31.4 percentage points from curation alone and an additional 10.3 points on average from post‑training.

By Chuxuan Hu, Hejie Cui, Norman Huang, Shubham Kumar Bharti, Wang-Chiew Tan, Sercan \"O. Ar{\i}k
arXiv Machine Learning
Sep 17

Behavior2Value: Benchmarking and Empowering LLMs for Consumer Value Measurement from E-commerce Behaviors

The paper introduces the Behavior-to-Value (B2V) task, which seeks to identify consumer values from e-commerce behavioral trajectories. It presents the E-commerce Consumption Value Taxonomy (ECVT) and the B2V-Bench dataset, derived from anonymized Taobao logs and covering 25 purchase behaviors with associated value orientations. A new model, B2V-Verifier, is proposed to improve value measurement accuracy, achieving a 34% boost in multi-label classification over strong LLM baselines.

By Peixuan Hou, Bin Chen, Li He, Jian Xu, Bo Zheng, Xiuli Ma, Guojie Song