INDRA is a research platform that integrates multiple archival collections—such as UCSF’s Industry Documents Library, Columbia and CUNY’s ToxicDocs, and Stanford’s SRITA—into a single, LLM‑readable corpus. It employs three safeguards: a closed evidentiary sandbox, real‑time provenance tagging, and a deterministic system‑level protocol to ensure that model outputs are clearly distinguished from archival evidence and from the model’s own inferences. The platform enables large‑language‑model‑powered investigations across these archives while keeping the conditions of knowledge production transparent and auditable.
By Daniel Akselrad, Robert N. Proctor
arXiv:2609.16519v1 Announce Type: new
Abstract: Scientific research increasingly relies on large, heterogeneous data sources, motivating interest in retrieval-augmented generation (RAG) systems that...
By Bernie Boscoe, Srinath Saikrishnan, Vikram Seenivasan, Jack Stark, Andrew Lizarraga, Morgan Himes, Jonathan Soriano, PJ Allen, Tuan Do
Scientific research increasingly relies on large, heterogeneous data sources, motivating interest in retrieval-augmented generation (RAG) systems that provide natural language access to scientific kno...
K‑Dense BYOK is a free, open‑source AI research assistant that runs locally on a researcher’s own computer. It provides a structured environment with scientific procedures, workflow templates, and a living lab notebook that logs all actions without allowing the agent to alter the record. The system emphasizes reproducibility by recording the software environment and offering commands to regenerate results, outperforming managed platforms on interdisciplinary research prompts.
By Aubrey M. Brueckner, Darshil Patel, Yuhuan He, Timothy Kassis
arXiv:2607. 11019v1 Announce Type: new Abstract: Enterprise data analysis is emerging as a distinct frontier for autonomous agents.
By Tianjing Zeng, Yuntao Hong, Zhongjun Ding, Dandan Liu, Yinan Mei, Yunxiang Su, Yiming Wang, Xiaojian Zhang, Jingyu Zhu, Junhao Zhu, Zhuowen Liang, Jiazhen Peng, Lianggui Weng, Zhihao Ding, Kerui Yi, Qifeng Wang, Rong Zhu, Bolin Ding, Liyu Mou, Jingren Zhou
arXiv:2607. 11464v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) addresses the limitations of Large Language Models (LLMs) when providing responses to domain-specific questions.
By Marlena Fl\"uh, Soo-Yon Kim, Carolin Victoria Schneider, Sandra Geisler
The paper introduces UniK, a universal knowledge perception platform designed to serve both digital AI—such as chatbots and agent workflows—and physical AI, which controls robots and autonomous systems. UniK handles the entire knowledge lifecycle—ingestion, enrichment, indexing, retrieval, and continuous evaluation—across diverse modalities including text, video, molecular data, and sensor telemetry, without task‑specific fine‑tuning. In five digital AI domains, UniK paired with a 70‑billion‑parameter model matches or surpasses larger proprietary LLMs, achieving high retrieval‑augmented generation accuracy on government data, medical QA, and chemistry tasks, and it also addresses similar data challenges in physical AI world‑model training.
By Nirmit Desai, Kunal Sawarkar, Aditya Mahakali, Dongkon Lee, Kevin Park, Eric Song
Nomad is an autonomous system designed to explore and discover insights within large data corpora. It builds an explicit Exploration Map to systematically traverse a domain, generating and testing hypotheses with an explorer agent that leverages document, web, and database searches. After verification, it produces cited reports and meta-reports, and its evaluation framework assesses trustworthiness, quality, and diversity, showing superior performance over baselines on UN, WHO, and arXiv datasets.
By Bokang Jia, Samta Kamboj, Satheesh Katipomu, Seung Hun Han, Neha Sengupta, Andrew Jackson
arXiv:2607. 28618v1 Announce Type: cross Abstract: Chemistry literature synthesis often requires assembling specific findings scattered across many publications, yet existing literature-search systems primarily return ranked document lists.
By Bing Yan, Gregory Wolfe, Stefano Martiniani, Kyunghyun Cho
arXiv:2608. 15193v1 Announce Type: cross Abstract: As large language model (LLM) agents are increasingly adopted in scientific research, external knowledge bases, knowledge graphs, and long-term memory have improved information retrieval and task continuity.
By Yuyang Zheng, Nan Li, Wenxia Deng, Lige Yan, Xiang Li, Si Chen
arXiv:2606. 26627v1 Announce Type: cross Abstract: Large language model agents increasingly query databases, search document collections, call external APIs, remember past interactions, and act on a user's behalf.
By Nada Lahjouji, Ashwin Gerard Colaco
arXiv:2606. 13669v1 Announce Type: new Abstract: Current LLM-based research agents have advanced through agent orchestration, yet largely overlook scientific knowledge orchestration.
By Zongsheng Cao, Bihao Zhan, Jinxin Shi, Jiong Wang, Fangchen Yu, Zhijie Zhong, Zijie Guo, Tianshuo Peng, Zhuo Liu, Yi Xie, Xiang Zhuang, Yue Fan, Runmin Ma, Shiyang Feng, Xiangchao Yan, Anran Liu, Peng Ye, Wenlong Zhang, Shufei Zhang, Chunfeng Song, Fenghua Ling, Jie Zhou, Liang He, Bo Zhang, Lei Bai