arXiv Machine Learning

Machine Learning for Coding Retail Product Names to Consumer-Price Categories: A Rule-plus-Bag-of-Words Pipeline with Reliability-Weighted Human-in-the-Loop Labeling

arXiv:2606. 02004v1 Announce Type: cross Abstract: Consumer-price measurement increasingly draws on alternative data sources -- scanner, web-scraped, and transaction/receipt data.

arXiv Computation and Language
Sep 1

LLP: LLM-Based Product Pricing in E-commerce

The paper introduces LLP, a Large Language Model–based generative framework for pricing second‑hand products on consumer‑to‑consumer platforms. LLP retrieves similar items to capture market dynamics, then uses LLMs to generate price suggestions, refined through supervised fine‑tuning and group relative policy optimization. A confidence‑based filter rejects unreliable predictions, and experiments show LLP outperforms prior methods, achieving higher static adoption rates when deployed on Xianyu.

By Hairu Wang, Sheng You, Qiheng Zhang, Xike Xie, Shuguang Han, Yuchen Wu, Fei Huang, Jufeng Chen
arXiv AI
Sep 1

Beyond Ranking Accuracy: Evaluating LLM-Cited Feature Rationales for Next Basket Repurchase Recommendation

arXiv:2608.30333v1 Announce Type: cross Abstract: Next-basket repurchase recommendation is commonly formulated as a ranking task: given a customer's purchase history, the system ranks previously purc...

By Yanan Cao, Anay Dombe, Murali Mohana Krishna Dandu, Shreeranjani Srirangamsridharan, Sinduja Subramaniam, Yogananth Mahalingam, Evren Korpeoglu, Kannan Achan
arXiv Machine Learning
Sep 17

Behavior2Value: Benchmarking and Empowering LLMs for Consumer Value Measurement from E-commerce Behaviors

The paper introduces the Behavior-to-Value (B2V) task, which seeks to identify consumer values from e-commerce behavioral trajectories. It presents the E-commerce Consumption Value Taxonomy (ECVT) and the B2V-Bench dataset, derived from anonymized Taobao logs and covering 25 purchase behaviors with associated value orientations. A new model, B2V-Verifier, is proposed to improve value measurement accuracy, achieving a 34% boost in multi-label classification over strong LLM baselines.

By Peixuan Hou, Bin Chen, Li He, Jian Xu, Bo Zheng, Xiuli Ma, Guojie Song
Hugging Face Trending Papers
Aug 18

Where A Small Language Model Helps in Invoice Categorisation, Understood Through Embedding Geometry

The paper explores using a small language model (SBERT) for invoice categorisation, a task that requires nuanced accounting judgement. By analysing the embedding geometry of SBERT and DeBERTa, the authors find that the sentence‑embedding space is globally anisotropic but contains locally isotropic clusters tied to vendor identity. Fine‑tuned SBERT achieves 0.96 accuracy and 0.9 F1 with only about 100 client‑specific invoices, outperforming zero‑shot LLMs and vendor baselines, and demonstrates that in‑house SLMs can reduce cost, enhance security, and improve interpretability.