The paper presents a two‑stage large language model pipeline for e‑commerce attribute extraction. First, it discovers a compact, ranked set of purchase‑discriminative attributes for each product category; second, it extracts those attribute values from catalog text using a fine‑tuned Qwen3‑4B model with Hyper‑Parallel Decoding. The approach attains 85% extraction accuracy while cutting inference costs by 92%, enabling scalable production use and producing consistent, comparable product knowledge bases.
By Nikhita Vedula, Dushyanta Dhyani, Bryan Wang, Shervin Malmasi
arXiv:2607. 07469v1 Announce Type: cross Abstract: Fine-tuning large language models (LLMs) for e-commerce attribute extraction requires labeled data representative across thousands of product types, attributes, and multiple languages.
By Andrea Scarinci, Virginia Negri, Brayan Impata, Suleiman Khan, Victor Martinez, Marcello Federico
arXiv:2609.15205v1 Announce Type: cross
Abstract: Table extraction from texts is an important task for information systems, and recent approaches that prompt large language models (LLMs) with instruc...
By Tong Li, Shuye Ding, Jiachuan Wang, Yongqi Zhang, Shuangyin Li, Lei Chen, Bo Li
The paper "Query Brand Entity Linking in E-Commerce Search" addresses the challenge of matching short, unstructured user search queries to the correct brand entities in large e‑commerce catalogs. It proposes two scalable solutions: a cascaded pipeline that first detects brand mentions and then disambiguates them, and a single‑stage extreme multiclass classifier that directly maps queries to brand identifiers. Extensive multilingual evaluation and an online experiment show that both methods significantly improve brand recall while preserving high precision, resulting in measurable gains in customer engagement.
By Dong Liu, Sreyashi Nag
arXiv:2609.07334v1 Announce Type: new
Abstract: The Asset Administration Shell (AAS) is a cornerstone of Industry 4.0 and the Digital Product Passport, providing standardized digital representations...
By Janek Gro{\ss}, Jens Heidrich
arXiv:2606. 04387v1 Announce Type: cross Abstract: Sales lead conversion in high-stakes domains (e.
By Chenyu Zhang, Yiwen Liu, Yin Sun, Xinyuan Zhang, Yuji Cao, Junming Jiao, Juyi Qiao
arXiv:2606. 26787v1 Announce Type: cross Abstract: Traditional dynamic pricing models in large-scale e-commerce suffer from limited interpretability, poor utilization of unstructured information, and misalignment with long-term business objectives such as cumulative Gross Merchandise Value (GMV), Return on Investment (ROI) and milestone achievement.
By Chennan Ma, Yanning Zhang, Siqi Hong, Xiuchong Wang, Fei Xiao, Keping Yang
The paper presents an iterative framework that uses large language models (LLMs) to automatically extract interpretable, schema‑bound categorical features from unstructured text for use in tabular prediction models. A generator LLM proposes semantic definitions, an extractor LLM materializes the features, and a downstream tabular model evaluates their predictive performance, with error‑driven natural‑language feedback guiding the search. Across three public datasets, the error‑driven loop speeds up feature discovery up to three times and the resulting features outperform any subset when combined with TF‑IDF and dense embeddings, while also providing instance‑level interpretability through SHAP importance rankings and a semantic audit trail.
By Merwan Barlier, Blaz Skrlj
Verifying the eligibility of securities as collateral is a key responsibility of the German Central Bank. However, manually verifying these assets against legal and financial criteria within lengthy, semi-structured, and often bilingual prospectuses is a resource-intensive task.
arXiv:2608. 06167v1 Announce Type: new Abstract: We present a schema-based framework for extracting complex, structured information from unstructured text documents using generative AI, followed by automated semantic evaluation of the extracted information against a gold standard.
By Modhurita Mitra, Jan-Willem Versteeg, Maarten D. Schermer, Shiva Nadi Najafabadi, Marie L. De Bruin, Lourens T. Bloem
LentEx is a new framework for latent entity extraction that uses synthetic data generation and instruction fine‑tuning to train smaller, efficient large language models. By creating diverse, contextually rich synthetic examples through a template‑based approach, LentEx overcomes the lack of labeled datasets and achieves strong performance, surpassing state‑of‑the‑art models on the MTEB Clustering Benchmark. The method also generalizes well to unseen domains, making it useful for tasks such as retrieval‑augmented generation, customer persona analysis, and knowledge graph enrichment.
By Umesh Bodhwani, Yuan Ling, Cibi Chakravarthy Senthilkumar, Shujing Dong, Yarong Feng, Hongfei Li, Ayush Goyal
arXiv:2502. 15411v4 Announce Type: replace-cross Abstract: Accurate tagging of earnings reports can yield significant short-term returns for stakeholders.
By Rasmus Aavang, Giovanni Rizzi, Rasmus B{\o}ggild, Alexandre Iolov, Mike Zhang, Johannes Bjerva