arXiv Machine Learning By Syed Mohammed Arshad Zaidi, Eric Rincon, Shayan Hassantabar

Serving the Long Tail: Training-Free LLM Candidate Generation for Vacation Rental Marketplaces

Read the original on arXiv Machine Learning →

arXiv:2607. 09877v1 Announce Type: new Abstract: Vacation rental marketplaces face a structural imbalance on the supply side: a small fraction of properties receive most user interactions, while the long tail of new, niche, and seasonal listings generates too little behavioral signal for collaborative filtering to serve effectively.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jun 30

Diagnosing and Mitigating Retrieval Bottlenecks in LLM-Based Cold-Start Recommendation

arXiv:2606. 29947v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as rerankers in recommender systems, with the expectation that semantic understanding will help in cold-start and long-tail regimes.

By Zhe Dong (University of Maine at Presque Isle), Fang Qin (Stanford University), Manish Shah (Independent Researcher), Yicheng Wang (Independent Researcher)
arXiv Machine Learning
Sep 10

A Multi-Source Ensemble Approach to Candidate Generation for Alternative Vacation Rental Property Recommendations

The paper studies candidate generation for alternative vacation rental recommendations, comparing collaborative filtering, shallow embeddings, and graph neural network (GNN) methods on a platform with over 2 million active properties. A hybrid model that combines item-based collaborative filtering with GNN-based retrieval achieves a 14.8% higher Recall@300 than the best baseline, leveraging each method’s strengths: collaborative filtering for well-interacted properties and GNNs for diverse, cold-start alternatives. The authors also show that stronger candidate pools improve downstream ranking quality, though the exact impact is intertwined with ranker training.

By Syed Mohammed Arshad Zaidi, Eric Rincon, Shayan Hassantabar
arXiv AI
Jun 10

STORM: Stepwise Token Optimization with Reward-Guided Beam Search

arXiv:2606. 10621v1 Announce Type: cross Abstract: Modern retrieval increasingly relies on dense and learned-sparse neural models that are effective but require encoding the entire corpus into a specialized index, rebuilt whenever the model changes.

By Arthur Satouf, Giulio D'Erasmo, Yuxuan Zong, Habiboulaye Amadou Boubacar, Pablo Piantanida, Benjamin Piwowarski
arXiv Machine Learning
1d ago

RPTune: Learned Context Curation for LLM Catalog Search

RPTune is an end‑to‑end framework that improves in‑context catalog search for small merchant businesses by learning to curate product catalogs and fine‑tuning large language models (LLMs) with catalog‑grounded supervision. It uses an encoder‑reorganizer curator to order and prune products based on LLM feedback, and then applies context‑relative rewards during LLM post‑training. Across seven real merchants and 100 complex conversational queries per merchant, RPTune boosts search accuracy by up to 31.4 percentage points from curation alone and an additional 10.3 points on average from post‑training.

By Chuxuan Hu, Hejie Cui, Norman Huang, Shubham Kumar Bharti, Wang-Chiew Tan, Sercan \"O. Ar{\i}k
arXiv Machine Learning
Jul 28

SMART: LLM-Augmented Hybrid Retrieval for Dynamic Product Ads

arXiv:2607. 23121v1 Announce Type: cross Abstract: Dynamic Product Ads (DPA) require retrieving relevant items from multi-million product catalogs, balancing two competing objectives: retargeting (re-surfacing known interests) and prospecting (discovering new categories).

By Congfei Zhang, Jingxiao Ma, Xiaodong Liu, Hsiang-wei Chao, Siman Wang, Ge Liu, Shantanu Aggarwal, Vincent Zhang, Meghana Missula, Rachel Liao, Zichu Li, Xiao Bai, Yunzhi Zhou, Yajun Wang, Zhe Liu, Jinchao Li, Yu Zhang