arXiv:2609.05760v1 Announce Type: cross
Abstract: We present RAGMark, a modular benchmarking framework for advanced Retrieval-Augmented Generation (RAG) systems targeting small-scale multi-GPU enviro...
By Zlatan Feric, Amir Taherin, Bin Ren, Yanzhi Wang, Jennifer Dy, David Kaeli
Orthrus is a serving system that performs heterogeneous batching of embedding and generative models within a single inference loop. It uses chunked embedding with incremental pooling and workload‑aware batch composition to unify conflicting computational patterns. Experiments on four A100 GPUs show that Orthrus improves throughput by 1.28×–4.52× and reduces p99 latency by up to 55.8% compared to baseline deployments.
By Dohyun Park, Hubertus Franke, Daniel G. Waddington, Swaminathan Sundararaman, Yongjoo Park
The paper introduces a GPU‑optimized retrieval framework for LinkedIn’s semantic search, partitioning embeddings into eight category‑supervised segments and applying a min/median aggregation rule aligned with the existing relevance policy. A lightweight Stage‑1 scorer generates high‑recall candidates, while a two‑stage GPU architecture—FP8 coarse ranking followed by FP16 re‑ranking—boosts throughput and recall, achieving 99.6‑99.8% of full‑FP16 recall at over 500 QPS per shard. In A/B testing, the system raises exploratory‑query Precision@10 from 63.7% to 79.0% and navigational Precision@1 from 65.5% to 74.7%, with human evaluation confirming the improvement.
By Dhritiman Das, Chujie Zheng, Ronak Kaoshik, Pratik Dixit, Vishal Shah, Yanbo Li, Jiahao Xu, Manika Agarwal, Chinmay Naik, Lingyu Zhang, Chetan Bhole, Chirag Bhanuprasad Mehta, Meng Zheng, Puneet Singh Ahluwalia, Shirisha Singh, Ping Jin, Manas Apte, Gokulraj Mohanasundaram, Tugrul Bingol, Raghavan Muthuregunathan, Fedor Borisyuk
arXiv:2607. 12392v1 Announce Type: cross Abstract: Optimizing large-scale retrieval hinges on the ability to efficiently surface candidates across diverse content tiers.
By Jiaxing Qu, Yilin Chen, Junpeng Hou, Jinfeng Rao, Olafur Gudmundsson, Sai Xiao, Huizhong Duan
KuaFu is a unified behavior‑compression layer that reduces each user behavior item to 2–4 tokens, dramatically shrinking per‑item cache size while preserving fidelity through a four‑stage training process. In production across four profiling tasks, it matches or outperforms uncompressed single‑task models, boosts GPU throughput by 37–350%, and saves 190 GPUs. On public benchmarks it consistently beats prior compressors at the same compression ratio, and on RecBench a 4B KuaFu model outperforms its 8B counterpart by 1.90 points, contributing to a 1.37% lift in overall GMV on Tencent’s advertising and recommendation platform.
By Jiahao Hui, Lin Zhu, Yishen Hu, Jingdong Shu, Zetai Jiang, Xining Ran, Ben Tan, Yeshou Cai, Gong Chen, Haijie Gu, Jie Jiang
arXiv:2602. 22647v2 Announce Type: replace-cross Abstract: Generative retrieval has emerged as a powerful paradigm for LLM-based recommendation.
By Zhengyang Su, Isay Katsman, Yueqi Wang, Ruining He, Lukasz Heldt, Raghunandan Keshavan, Shao-Chuan Wang, Xinyang Yi, Mingyan Gao, Onkar Dalal, Lichan Hong, Ed Chi, Ningren Han
The paper introduces Connected Content Retriever (CC Retriever), a pre‑ranking system for LinkedIn’s Feed that uses dense graph edge features to score candidate content from a billion‑scale index within a 120 ms latency budget. By leveraging GPU‑based sorted‑search primitives, the system can apply a full deep ranking model with 50× more parameters, achieving a 2.5% lift in content time spent in online experiments. The work details the economic‑graph features and model architecture that enable this scalable, low‑latency scoring pipeline.
By Akhilesh Gupta, Sudarshan Srinivasa Ramanujam, Chirag Bhanuprasad Mehta, Reshma Asharaf Beena, Dhritiman Das, Birjodh Singh Tiwana, Bhargavkumar Kanubhai Patel, Mack Lee, Renyi Tang
arXiv:2607. 10044v1 Announce Type: new Abstract: Constrained decoding is essential in generative retrieval, where document identifiers generated directly from a query must exactly match a predefined library of valid IDs.
By Dakshitha Anandakumar, Anurag Mukkara, Wenxiang Hu, Jiusheng Chen, M Akash Kumar, Ting Ye, Qiang Lou, Jian Jiao
arXiv:2607. 27090v1 Announce Type: cross Abstract: Large language models are increasingly deployed with persistent personalized context, such as accumulated memory profiles or long conversation histories, that is shared across a user's many requests.
By Peter Li, Prashant Pandey
arXiv:2609.38090v1 Announce Type: new
Abstract: Mixture-of-Experts (MoE) models are a compelling architecture for scaling model capacity, making them especially attractive for deployment on resource-...
By Sanjali Yadav, Bahar Asgari
arXiv:2604. 26968v2 Announce Type: replace-cross Abstract: Key-value (KV) cache memory management is the primary bottleneck limiting throughput and cost-efficiency in large-scale GPU inference serving.
By Sanjeev Rao Ganjihal
Large language models are increasingly deployed with persistent personalized context, such as accumulated memory profiles or long conversation histories, that is shared across a user's many requests. Production memory systems (e.