Using a Large Language Model (LLM) as the clusterer at production scale is hard: prompts cannot hold the entire label space, and per-document serial processing does not deliver the throughput real wor...
arXiv:2606. 31156v1 Announce Type: cross Abstract: RAG systems retrieve documents optimized for answering one query at a time.
By Shivam Ratnakar, Yixuan Zhu, Cecilia Cheng, Chaya Vijayakumar
arXiv:2607. 19704v1 Announce Type: new Abstract: Scaling LLM-based applications to millions of users is bottlenecked by the inference cost and latency of modern foundation models.
By Longshaokan Wang, Wai Tsang Keung, Punit Ghodasara, Roman Wang, Ali Dashti, Francesc Moreno-Noguer
The paper proposes an incremental pooled LLM evaluation method for selecting retrieval models in production Retrieval-Augmented Generation (RAG) systems. By having a language model judge the union of documents retrieved by current candidates and expanding the pool only with new documents from added systems, the approach reuses judgments across all systems. Experiments on four benchmarks and a financial news QA deployment show strong correlation with gold-standard rankings, high preservation of pairwise orderings, and significant cost savings—up to 4.9× lower evaluation cost and 65–80% judgment reuse.
By Max Nelson, Hanoz Bhathena, Aviral Joshi, Saket Sharma
arXiv:2609.10239v1 Announce Type: cross
Abstract: Graph-based retrieval can improve multi-hop question answering, but existing approaches often incur high query-time costs and produce diffuse, oversi...
By Daniel Alejandro Coll Tejeda, Pedro Garc\'ia L\'opez, Daniel Barcelona-Pons
arXiv:2606. 29947v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as rerankers in recommender systems, with the expectation that semantic understanding will help in cold-start and long-tail regimes.
By Zhe Dong (University of Maine at Presque Isle), Fang Qin (Stanford University), Manish Shah (Independent Researcher), Yicheng Wang (Independent Researcher)
arXiv:2606. 28328v1 Announce Type: cross Abstract: In recent years, text clustering has become a critical technique for applications including intent discovery, topic mining, and recommendation systems.
By Daoming Wan, Yizheng Huang, Jimmy X. Huang
The paper introduces ONLINE LLM PICKER, a framework for active model selection of large language models in streaming settings. It selects the most informative prompts for annotation within a limited budget, enabling the identification of the best or near‑best model among many candidates. Experiments on 10 datasets and over 130 language models show up to 71.67% savings in annotation cost and a reduction in regret by up to 2.51× when using the chosen model for sequential generation.
By Alessandro Turrin, Patrik Okanovic, Torsten Hoefler, Nezihe Merve G\"urel
arXiv:2607. 10548v1 Announce Type: cross Abstract: Pseudo-labeling based on Optimal Transport (OT) has become an effective mechanism for enhancing short text clustering.
By Zhihao Yao, Yuxuan Gu, Jixuan Yin, Bo Li
arXiv:2603. 26815v3 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) systems for financial document QA typically follow a chunk-based paradigm: documents are split into fragments, embedded, and retrieved by similarity.
By Zhiyuan Cheng, Longying Lai, Yue Liu
arXiv:2607. 12392v1 Announce Type: cross Abstract: Optimizing large-scale retrieval hinges on the ability to efficiently surface candidates across diverse content tiers.
By Jiaxing Qu, Yilin Chen, Junpeng Hou, Jinfeng Rao, Olafur Gudmundsson, Sai Xiao, Huizhong Duan
arXiv:2511. 16681v3 Announce Type: replace-cross Abstract: Vector databases (VecDBs) are increasingly deployed in retrieval-augmented generation (RAG) pipelines where query processing and document ingestion occur concurrently.
By Dong Liu, Yanxuan Yu