arXiv AI

KadiAssistant: A conversational AI Agent for information retrieval in Kadi4Mat

arXiv Computation and Language
Sep 11

INDRA: A New AI Tool for Exploring Tobacco, Fossil Fuel, and Chemical Industry Archives

INDRA is a research platform that integrates multiple archival collections—such as UCSF’s Industry Documents Library, Columbia and CUNY’s ToxicDocs, and Stanford’s SRITA—into a single, LLM‑readable corpus. It employs three safeguards: a closed evidentiary sandbox, real‑time provenance tagging, and a deterministic system‑level protocol to ensure that model outputs are clearly distinguished from archival evidence and from the model’s own inferences. The platform enables large‑language‑model‑powered investigations across these archives while keeping the conditions of knowledge production transparent and auditable.

By Daniel Akselrad, Robert N. Proctor
arXiv AI
Sep 16

AquiLLM: Evaluating Faithfulness in Open-Weight RAG-LLM Systems for Scientific Research

arXiv:2609.16519v1 Announce Type: new Abstract: Scientific research increasingly relies on large, heterogeneous data sources, motivating interest in retrieval-augmented generation (RAG) systems that...

By Bernie Boscoe, Srinath Saikrishnan, Vikram Seenivasan, Jack Stark, Andrew Lizarraga, Morgan Himes, Jonathan Soriano, PJ Allen, Tuan Do
arXiv AI
6d ago

K-Dense BYOK: An Open-Source AI Research Assistant That Runs Locally and Keeps a Hash-Chained Lab Notebook

K‑Dense BYOK is a free, open‑source AI research assistant that runs locally on a researcher’s own computer. It provides a structured environment with scientific procedures, workflow templates, and a living lab notebook that logs all actions without allowing the agent to alter the record. The system emphasizes reproducibility by recording the software environment and offering commands to regenerate results, outperforming managed platforms on interdisciplinary research prompts.

By Aubrey M. Brueckner, Darshil Patel, Yuhuan He, Timothy Kassis
arXiv AI
Jul 14

QwenPaw-Data: Bridging Facts, Methodology, and Execution for Autonomous Enterprise Data Analytics

arXiv:2607. 11019v1 Announce Type: new Abstract: Enterprise data analysis is emerging as a distinct frontier for autonomous agents.

By Tianjing Zeng, Yuntao Hong, Zhongjun Ding, Dandan Liu, Yinan Mei, Yunxiang Su, Yiming Wang, Xiaojian Zhang, Jingyu Zhu, Junhao Zhu, Zhuowen Liang, Jiazhen Peng, Lianggui Weng, Zhihao Ding, Kerui Yi, Qifeng Wang, Rong Zhu, Bolin Ding, Liyu Mou, Jingren Zhou
arXiv Machine Learning
Sep 22

UniK: Universal Knowledge Perception for Digital and Physical AI

The paper introduces UniK, a universal knowledge perception platform designed to serve both digital AI—such as chatbots and agent workflows—and physical AI, which controls robots and autonomous systems. UniK handles the entire knowledge lifecycle—ingestion, enrichment, indexing, retrieval, and continuous evaluation—across diverse modalities including text, video, molecular data, and sensor telemetry, without task‑specific fine‑tuning. In five digital AI domains, UniK paired with a 70‑billion‑parameter model matches or surpasses larger proprietary LLMs, achieving high retrieval‑augmented generation accuracy on government data, medical QA, and chemistry tasks, and it also addresses similar data challenges in physical AI world‑model training.

By Nirmit Desai, Kunal Sawarkar, Aditya Mahakali, Dongkon Lee, Kevin Park, Eric Song
arXiv AI
Aug 28

Nomad: Autonomous Exploration and Discovery

Nomad is an autonomous system designed to explore and discover insights within large data corpora. It builds an explicit Exploration Map to systematically traverse a domain, generating and testing hypotheses with an explorer agent that leverages document, web, and database searches. After verification, it produces cited reports and meta-reports, and its evaluation framework assesses trustworthiness, quality, and diversity, showing superior performance over baselines on UN, WHO, and arXiv datasets.

By Bokang Jia, Samta Kamboj, Satheesh Katipomu, Seung Hun Han, Neha Sengupta, Andrew Jackson
arXiv AI
Jun 12

Agents-K1: Towards Agent-native Knowledge Orchestration

arXiv:2606. 13669v1 Announce Type: new Abstract: Current LLM-based research agents have advanced through agent orchestration, yet largely overlook scientific knowledge orchestration.

By Zongsheng Cao, Bihao Zhan, Jinxin Shi, Jiong Wang, Fangchen Yu, Zhijie Zhong, Zijie Guo, Tianshuo Peng, Zhuo Liu, Yi Xie, Xiang Zhuang, Yue Fan, Runmin Ma, Shiyang Feng, Xiangchao Yan, Anran Liu, Peng Ye, Wenlong Zhang, Shufei Zhang, Chunfeng Song, Fenghua Ling, Jie Zhou, Liang He, Bo Zhang, Lei Bai