arXiv Computation and Language By Nikhita Vedula, Dushyanta Dhyani, Bryan Wang, Shervin Malmasi

Scaling E-Commerce Attribute Extraction with Parallel Decoding

Read the original on arXiv Computation and Language →

The paper presents a two‑stage large language model pipeline for e‑commerce attribute extraction. First, it discovers a compact, ranked set of purchase‑discriminative attributes for each product category; second, it extracts those attribute values from catalog text using a fine‑tuned Qwen3‑4B model with Hyper‑Parallel Decoding. The approach attains 85% extraction accuracy while cutting inference costs by 92%, enabling scalable production use and producing consistent, comparable product knowledge bases.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Jun 24

AI-PAVE-Br: Leveraging Large Language Models for Enhanced Product Attribute Value Extraction through a Golden Set Approach

arXiv:2606. 24655v1 Announce Type: cross Abstract: The explosive growth and complexity of product data within the dynamic Brazilian e-commerce landscape demand robust and specialized methods for structured information extraction.

By Murilo Gazzola, Hugo Gobato Souto, Samuel Silva, J\'ulia Schubert Peixoto, Felipe Siqueira, Andr\'e Luis Pedroso de Morais, Caio Gomes
arXiv Machine Learning
Jun 2

Semantic Retrieval for Product Search in E-Commerce

arXiv:2606. 01504v1 Announce Type: cross Abstract: Semantic retrieval in e-commerce must handle short, noisy, and colloquial queries over large product catalogs with fine-grained attribute distinctions.

By Nikhil Kothari, Saksham Samdani, Ritam Mallick, Praveen Gupta, Ankit Vijay, Surender Kumar
arXiv Computation and Language
6d ago

REALMS: An AI-Assistant Conversational System for Real-Time Exact Audience Sizing over High-Dimensional Nested Profiles

REALMS is a conversational system that provides real‑time, exact audience sizing for digital marketers. It uses embedding‑based vector search to retrieve relevant categorical attributes, an LLM‑powered NL2SQL pipeline for accurate query generation over complex nested schemas, and schema standardization for industry‑agnostic deployment. Evaluations on real enterprise data show high recall, accurate SQL execution, and low latency, enabling interactive audience insights that previously required hours.

By Haixu Ma, Aditya Bansal, Shubham Lohiya, Sumit Ranjan
arXiv AI
Aug 24

One Hierarchy, Two Systems: Semantic Product IDs for Discovery-Surface Ranking and Search-Page Query Reformulation

The paper proposes a single hierarchical Semantic ID (SID) system to unify product identification across multiple merchants in e-commerce. By learning SID representations from product content, the authors demonstrate that ranking algorithms can aggregate consumer affinity and product performance over SID prefixes, improving offline relevance and online engagement. For query reformulation, SID concepts guide navigation and refinement, yielding better intent preservation and higher-quality suggestions compared to taxonomy or raw query transitions.

By Steven Xu, Sanjyot Thete, Saathvik Dirisala, Raghav Saboo, Nimesh Sinha, Leo Shao, Elyse Winer, Sudeep Das, Martin Wang, Kyle MacDonald