arXiv AI By Jack Stark, Srinath Saikrishnan, Vikram Seenivasan, Bernie Boscoe, Andrew Lizarraga, Tuan Do

AquiLLM: An Architecture for Supporting Tacit Knowledge Capture in Research Groups

Read the original on arXiv AI →

arXiv:2608. 08883v1 Announce Type: new Abstract: Recent advances in retrieval-augmented generation (RAG) and large language models (LLMs) enable researchers to integrate AI into scientific workflows.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 16

AquiLLM: Evaluating Faithfulness in Open-Weight RAG-LLM Systems for Scientific Research

arXiv:2609.16519v1 Announce Type: new Abstract: Scientific research increasingly relies on large, heterogeneous data sources, motivating interest in retrieval-augmented generation (RAG) systems that...

By Bernie Boscoe, Srinath Saikrishnan, Vikram Seenivasan, Jack Stark, Andrew Lizarraga, Morgan Himes, Jonathan Soriano, PJ Allen, Tuan Do
arXiv AI
2d ago

Mapping the RAG Landscape: A Four Axis Taxonomy of Efficiency, Defense, Interactivity, and Reasoning

The paper surveys recent advances in Retrieval Augmented Generation (RAG), a technique that integrates external retrieval into language model generation to reduce hallucinations and keep knowledge current. It introduces a four‑axis taxonomy—efficiency, defense, interactivity, and reasoning—to organize contemporary RAG research, covering retrieval methods, fusion strategies, embedding optimizations, and reinforcement learning policies. The survey also reviews evaluation practices, domain‑specific applications, and architectural variants, while highlighting ongoing challenges such as retrieval quality, reliability, domain adaptation, scalability, and explainability.

By Meghana Sunil, Shravya V, Shravan Venkatraman, Joe Dhanith PR
arXiv Machine Learning
Sep 18

A Large-Scale Vision-Language Dataset Derived from Open Scientific Literature to Advance Biomedical Generalist AI

The paper introduces Biomedica, an open-source dataset sourced from PubMed Central that includes over 6 million scientific articles and 24 million image‑text pairs, along with 27 metadata fields and expert human annotations. To facilitate use, the authors provide scalable streaming and search APIs via a web server. They demonstrate the dataset’s value by training embedding models, chat‑style models, and retrieval‑augmented chat agents, all of which outperform previous open systems in their categories.

By Alejandro Lozano, Min Woo Sun, James Burgess, Jeffrey J. Nirschl, Christopher Polzak, Yuhui Zhang, Liangyu Chen, Jeffrey Gu, Ivan Lopez, Josiah Aklilu, Anita Rau, Austin Wolfgang Katzer, Collin Chiu, Orr Zohar, Xiaohan Wang, Alfred Seunghoon Song, Chiang Chia-Chun, Robert Tibshirani, Serena Yeung-Levy