Retrieval-augmented generation

Retrieval pipelines, vector search, chunking and reranking: how models are grounded in a corpus instead of their weights.

3,469 stories · RSS feed

arXiv Machine Learning
Jun 3

Explainable Forecasting of Scientific Breakthroughs from Concept Network Dynamics

arXiv:2606. 03864v1 Announce Type: cross Abstract: We introduce an explainable machine-learning approach that forecasts the structural precursors of scientific breakthroughs -- the emergence and intensification of links between research concepts -- by modelling how OpenAlex concept networks evolve over time.

By Thomas Maillart, Thibaut Chataing, Ntorina Antoni, David Dosu, Paul Bagourd, Julian Jang-Jaccard, Alain Mermoud
arXiv Machine Learning
Jun 3

The Efficiency vs. Accuracy Trade-off: Optimizing RAG-Enhanced LLM Recommender Systems Using Multi-Head Early Exit

arXiv:2501. 02173v2 Announce Type: replace-cross Abstract: The deployment of Large Language Models (LLMs) in recommender systems for predicting Click-Through Rates (CTR) necessitates a delicate balance between computational efficiency and predictive accuracy.

By Huixue Zhou, Hengrui Gu, Xi Liu, Kaixiong Zhou, Mingfu Liang, Yongkang Xiao, Srinivas Govindan, Piyush Chawla, Jiyan Yang, Xiangfei Meng, Huayu Li, Buyun Zhang, Liang Luo, Wen-Yen Chen, Yiping Han, Bo Long, Rui Zhang, Tianlong Chen
arXiv AI
Jun 3

vLLM Semantic Router: Signal Driven Decision Routing for Mixture-of-Modality Models

arXiv:2603. 04444v3 Announce Type: replace-cross Abstract: As large language models (LLMs) diversify across modalities, capabilities, and cost profiles, the problem of intelligent request routing -- selecting the right model for each query at inference time -- has become a critical systems challenge.

By Xunzhuo Liu (Steve), Huamin Chen (Steve), Samzong Lu (Steve), Yossi Ovadia (Steve), Guohong Wen (Steve), Hao Wu (Steve), Zhengda Tan (Steve), Jintao Zhang (Steve), Senan Zedan (Steve), Yehudit Kerido (Steve), Liav Weiss (Steve), Haichen Zhang (Steve), Bishen Yu (Steve), Asaad Balum (Steve), Noa Limoy (Steve), Abdallah Samara (Steve), Baofa Fan (Steve), Brent Salisbury (Steve), Ryan Cook (Steve), Zhijie Wang (Steve), Qiping Pan (Steve), Rehan Khan (Steve), Avishek Goswami (Steve), Houston H. Zhang (Steve), Shuyi Wang (Steve), Ziang Tang (Steve), Fang Han (Steve), Zohaib Hassan (Steve), Jianqiao Zheng (Steve), Avinash Changrani (Steve), Xue (Steve), Liu, Bowei He
arXiv Machine Learning
Jun 3

Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning

arXiv:2603. 01471v3 Announce Type: replace-cross Abstract: Multimodal embedding models, rooted in multimodal large language models (MLLMs), have yielded significant performance improvements across diverse tasks such as retrieval and classification.

By Jiahan Chen, Da Li, Hengran Zhang, Yinqiong Cai, Lixin Su, Jiafeng Guo, Daiting Shi, Dawei Yin, Keping Bi
arXiv AI
Jun 3

GFFMERGE: Efficient Merging of Graph Neural Force Fields and Beyond

arXiv:2606. 03232v1 Announce Type: cross Abstract: Graph Neural Networks (GNNs) have revolutionized Neural Force Fields for atomistic simulations, achieving near-quantum accuracy at reduced cost, yet adapting these models to new chemical systems requires expensive retraining of foundation models.

By Parth Verma, Parv P. Singh, Vipul Garg, Ishita Thakre, N. M. Anoop Krishnan, Sayan Ranu
arXiv AI
Jun 3

UR$^2$: Unify RAG and Reasoning through Reinforcement Learning

arXiv:2508. 06165v5 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have shown strong capabilities through two complementary paradigms: Retrieval-Augmented Generation (RAG) for knowledge grounding and Reinforcement Learning from Verifiable Rewards (RLVR) for complex reasoning.

By Weitao Li, Boran Xiang, Xiaolong Wang, Zhinan Gou, Weizhi Ma, Yang Liu
arXiv AI
Jun 3

Inference Cost Attacks for Retrieval-Augmented Large Language Models

arXiv:2606. 02643v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG)-enhanced LLM systems, while powerful, introduce substantial inference costs due to the inclusion of an extra multi-stage pipeline that dynamically retrieves and synthesizes information from external knowledge sources.

By Chengliang Liu, Liangbo Ning, Yujuan Ding, Wenqi Fan