arXiv AI By Ivan Ji (Zihao), Liuyi Hu (Zihao), Harrison (Zihao), Zhao (Xiangjun), Lei Huang (Xiangjun), Qunshu Zhang (Xiangjun), Max (Xiangjun), Fan, Aameek Singh

Real-Time Hard Negative Sampling via LLM-based Clustering for Large-Scale Two-Tower Retrieval

Read the original on arXiv AI →

arXiv:2607. 00448v1 Announce Type: cross Abstract: The two-tower model has been widely used for large-scale recommendation systems, particularly in the retrieval stage.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 5

From Generator to Embedder: Harnessing Innate Abilities of Multimodal LLMs via Building Zero-Shot Discriminative Embedding Model

arXiv:2508. 00955v3 Announce Type: replace-cross Abstract: Adapting generative Multimodal Large Language Models (MLLMs) into universal embedding models typically demands resource-intensive contrastive pre-training, while traditional hard negative mining methods suffer from severe false negative contamination.

By Yeong-Joon Ju, Seong-Whan Lee
arXiv Computation and Language
Sep 11

LLMAR: A Tuning-Free Recommendation Framework for Sparse and Text-Rich Industrial Domains

LLMAR is a tuning‑free recommendation framework designed for sparse, text‑rich industrial B2B domains. It transforms user behavioral history into structured semantic motives using LLM inference, employs a reflection loop to self‑correct hallucinations, and operates cost‑effectively with asynchronous batch processing. Experiments on MovieLens‑1M, Amazon Prime Pantry, and a construction risk dataset show LLMAR surpasses state‑of‑the‑art learning models, achieving up to a 54.6% nDCG@10 improvement while keeping inference costs around $1 per 1,000 users.

By Ryogo Hishikawa, Ichiro Kataoka, Shinya Yuda
arXiv Machine Learning
Aug 28

Recipes for Steering and Scaling LLMs via Sampling

The paper introduces a flexible, theoretically grounded framework for steering and scaling autoregressive large language models (LLMs) through sampling. It presents two algorithms—Sequential Monte Carlo (SMC) and Replica Exchange (RE)—that guide generation toward desired distributions such as powering, product, or tilting of the base model. Experiments show these methods outperform Best‑of‑N and standard MCMC baselines, offering a systematic recipe for probabilistic inference with LLMs via sampling.

By Jiajun He, Zongyu Guo, Jos\'e Miguel Hern\'andez-Lobato, Yuanqi Du