arXiv:2606. 29328v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) typically treats context selection as ranking chunks against a single query embedding.
By Bingxue Zhang, Jianying Jia, Feida Zhu
arXiv:2606. 15216v1 Announce Type: cross Abstract: Diversity plays a critical role in data selection, improving performance under fixed data budgets by reducing redundancy and repetition.
By Clarence Lee, Yejin Choi, Luke Zettlemoyer, Pang Wei Koh, Hai Leong Chieu
arXiv:2607. 11684v1 Announce Type: cross Abstract: Existing contextual multinomial logit (MNL) bandits model relevance-driven choice but ignore the potential benefits of within-assortment diversity, while submodular/combinatorial bandits encode diversity in rewards but lack structured choice probabilities.
By Heesang Ann, Taehyun Hwang, Min-hwan Oh
arXiv:2607. 09739v1 Announce Type: new Abstract: We study LLM benchmark coreset selection: selecting a small subset of prompts over multiple benchmarks whose induced model scores and rankings approximate those obtained from the full benchmark suite.
By Jihan Yao, Gantavya Bhatt, Arnav Das, Peter Jin, Ke Bao, Qiaolin Yu, Khushi Bhardwaj, Chang Su, Jialei Wang, Yikai Zhu, Sugam Devare, Damon Mosk-Aoyama, Zhen Dong, Venkat Krishna Srinivasan, Yineng Zhang, Oleksii Kuchaiev, Jiantao Jiao, Banghua Zhu, Jeff Bilmes
arXiv:2510. 16882v4 Announce Type: replace-cross Abstract: Supervised fine-tuning (SFT) is a commonly used technique to adapt large language models (LLMs) to downstream tasks.
By Heming Zou, Yixiu Mao, Yun Qu, Qi Wang, Xiangyang Ji
The paper examines LLM-based recommendation rerankers that are often evaluated under an oracle protocol, which guarantees the ground-truth item is present in the scored set. Across Amazon datasets, this protocol overestimates realistic NDCG@10 by 92–95% because realistic retrieval only covers 2–19% of relevant items at K=100, creating a recall ceiling that limits any closed-candidate reranker's top‑k NDCG. The authors find that various optimisation strategies—including prompt engineering, model scaling, sequential models, supervised neural rerankers, LoRA fine‑tuning, hybrid retrieval, score‑aware prompting, and LLM+CF fusion—do not significantly improve over a collaborative‑filtering baseline under realistic retrieval, and they propose a Recall‑Aware Evaluation Protocol (RAEP) to better assess rerankers in low‑recall regimes.
By Zhaohui Wang