Rethinking Sales Lead Scoring with LLM-based Hierarchical Preference Ranking
arXiv:2606. 04387v1 Announce Type: cross Abstract: Sales lead conversion in high-stakes domains (e.
arXiv:2606. 04387v1 Announce Type: cross Abstract: Sales lead conversion in high-stakes domains (e.
arXiv:2606. 02004v1 Announce Type: cross Abstract: Consumer-price measurement increasingly draws on alternative data sources -- scanner, web-scraped, and transaction/receipt data.
arXiv:2607. 23647v1 Announce Type: cross Abstract: Large language models (LLMs) can summarize heterogeneous user evidence in natural language, but current LLM recommenders often collapse enduring preferences, transient intent, and exposure-induced behavior into one profile.
The paper proposes replacing multiple horizon‑specific binary classifiers with a single survival model to predict time‑to‑repurchase in grocery e‑commerce. Empirical analysis shows a slightly decreasing hazard (k≈0.9) and that a Log‑Normal model best fits marginal distributions while Weibull best fits residuals. A single Accelerated Failure Time (AFT) model matches or surpasses per‑horizon classifiers with fewer trees, and a 4‑parameter calibration maps survival CDFs to horizon probabilities without monotonicity violations, revealing a trade‑off between calibration and ranking within the AFT family.
arXiv:2607. 06993v1 Announce Type: new Abstract: Customer behavior modeling underpins recommendation, marketing, and decision support, yet existing approaches either optimize predictive accuracy without explaining decisions or simulate users without grounding them in real behavioral data.
arXiv:2606. 06779v1 Announce Type: cross Abstract: In multi-vertical e-commerce platforms like DoorDash, relatively newer product verticals such as grocery and retail present a significant opportunity for personalization innovation.
The paper introduces DCEO, a data‑driven framework that learns item‑level proxy scores directly aligned with long‑term user objectives in e‑commerce search. It aggregates these scores into a user‑level metric, measures alignment via relative causal effect, and uses an actor‑critic model to generate context‑dependent fusion weights for multiple objectives. Offline experiments and a 41‑day online A/B test show DCEO improves GMV by 0.36% over traditional proxies.
The paper investigates how large language models (LLMs) can evaluate explanations in recommender systems. It generates 18 explanation prototypes and has 14 LLMs rate them, comparing the results to human ratings from a user study. Findings show that while LLMs mimic human rating patterns and correlate moderately with human judgments, their absolute agreement is low and varies with model size and evaluation design, leading to four practical recommendations for using LLMs in this context.
arXiv:2606. 23701v1 Announce Type: cross Abstract: Qualitative product feedback can reveal nuanced user experiences, but its implicit sentiment is difficult to measure.
arXiv:2608.20801v1 Announce Type: cross Abstract: While Large Language Models (LLMs) have significantly advanced reranking in recommendation, effectively leveraging item-side information remains chal...
arXiv:2607. 25420v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used in recommender systems, but it is often unclear how much performance can be obtained from strong pre-trained backbones alone when they are placed inside a structured recommendation pipeline.
arXiv:2603. 29247v3 Announce Type: replace-cross Abstract: LLM-based shopping agents increasingly rely on long purchase histories and multi-turn interactions for personalization, yet naively appending raw history to prompts is often ineffective due to noise, length, and relevance mismatch.