arXiv AI

K-Dense BYOK: An Open-Source AI Research Assistant That Runs Locally and Keeps a Hash-Chained Lab Notebook

K‑Dense BYOK is a free, open‑source AI research assistant that runs locally on a researcher’s own computer. It provides a structured environment with scientific procedures, workflow templates, and a living lab notebook that logs all actions without allowing the agent to alter the record. The system emphasizes reproducibility by recording the software environment and offering commands to regenerate results, outperforming managed platforms on interdisciplinary research prompts.

arXiv AI
Sep 1

CrossAudit: A Git-Native, Cross-Vendor Audit Loop for Agentic Science

CrossAudit proposes a Git‑native, cross‑vendor audit protocol for autonomous research pipelines, ensuring each work increment is reviewed by an agent from a different vendor against a human‑written rulebook. Audit outcomes, disputes, and rulings are stored as git commits, providing a replayable, versioned supervision history. The authors implemented the protocol with GitHub Actions and Python, deployed it in a computational‑chemistry pipeline, and conducted a seeded‑defect trial that revealed differing interpretations of the same rulebook by two vendors.

By Zhaohe Dong, Yuhao Chen
arXiv AI
Aug 19

StagedWorkspace: A Versioned Workspace for Knowledge-Work Agents

StagedWorkspace is a versioned workspace designed for knowledge‑work AI agents, ensuring that every parsed view, native file edit, and review diff is explicitly tied to a specific version of the workspace state. By binding parsed records and review diffs to content hashes of native files, the system improves performance on tasks such as OfficeQA and APEX‑Agents, achieving higher pass rates and rubric scores compared to single‑view approaches. The study demonstrates that providing dual parsed/native access and visible diffs enhances agent performance, highlighting workspace state as a key experimental variable for future benchmarks.

By Yining Hua, Hongbin Na, Yifan Zhou, Akshay Kalose, Cyrus Ayubcha, Levi Lian
arXiv AI
Sep 2

Dr. Claw: An AI Scientist Workspace for Vibe Research

Dr. Claw is an open‑source AI scientist workspace that integrates existing command‑line coding agents into a single, auditable, human‑in‑the‑loop workflow. It uses persistent state objects, a reusable skill library, and multi‑executor coordination to link human decisions with AI execution, creating a traceable and recoverable loop for planning, execution, and writing. The authors demonstrate the system with an interactive scenario and a failure‑recovery walkthrough, and show that, when the underlying executor is held constant, Dr. Claw achieves higher research completeness while preserving an auditable process trail.

By Dingjie Song, Hanrong Zhang, Dawei Liu, Yixin Liu, Zongxia Li, Zhengqing Yuan, Siqi Zhang, Henry Peng Zou, Zhiling Yan, Yuxuan Zhang, Yanfang Ye, Philip S. Yu, Lichao Sun
arXiv AI
Sep 25

Low-Cost Assays for Measuring Model Behavior Across Vendors and Releases

The paper introduces inexpensive, scalable methods for evaluating language model behavior across different vendors and releases. By running identical public stimuli on a cross‑vendor panel and analyzing transcripts via exact match, LLM‑coded codebooks, or instrumented environments, the authors can quantify model responses at a cost of a few dollars per model. Applying these tools to four years of releases reveals patterns of convergence, resistance, positional stability, and compliance that vary by generation, lab, and harness.

By Tapan Parikh
arXiv AI
Sep 3

Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills

The paper introduces Repo-To-Skill, a method for converting GitHub repositories into reusable AI skills. By distilling operational knowledge from over 1,000 machine‑learning repositories, the authors build the AREX‑Skill Library with more than 5,000 verified skills across 20 areas. Integrating these skills into a research agent—DisCo—yields significant performance boosts on multiple benchmarks, demonstrating the value of reusable, task‑agnostic knowledge.

By Jianlyu Chen, Yuyang Hu, Hongjin Qian, Jiawei Liu, Wenqing Wei, Xiaolong Chen, Defu Lian, Zhicheng Dou, Chaozhuo Li, Qiwei Ye, Zheng Liu
arXiv AI
Sep 7

TruthInsightBench: An Evidence-Grounded Benchmark for Automated Evaluation of Open-Ended Scientific Discovery Agents

TruthInsightBench is a new benchmark designed to evaluate automated scientific discovery agents by presenting them with 40 blind tasks drawn from peer‑reviewed studies across ten domains. Each task provides only a neutral objective and frozen data, withholding source conclusions, expected values, and analysis paths, forcing agents to determine which claim the data support. A fixed LLM‑based judge scores agents on evidentiary maturity across six dimensions, using 29 artifact‑grounded items, enabling fully automated, repeatable evaluation without human grading.

By Zhibo Yang, Chen Zhang, Yuewei Zhang, Hao Wang
arXiv AI
2d ago

YouRA: A Persistent-State Architecture for Evidence-Traceable Autonomous Research Agents

YouRA (Your Research Agent) is a persistent-state architecture that enables evidence‑traceable autonomous research agents to maintain research state, execution evidence, and failure history across long‑horizon pipelines. It combines a Verification State Architecture to track hypotheses and evidence, an Independent Controller to manage lifecycle and recovery, and Stateful Reflection to log failures and guide repair. On the MLR‑Bench ten‑task subset, YouRA outperforms MLR‑Agent and AI Scientist V2, and ablation studies confirm the importance of each component.

By Yoonkyu Woo, Woojin Lee, Jin-Xia Huang
Hugging Face Trending Papers
Sep 2

Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills

Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills introduces DisCo, a research agent that extracts and verifies operational knowledge from GitHub repositories to create reusable AI skills. The agent produces both task‑agnostic skills—compiled into the AREX‑Skill Library of over 5,000 verified skills from 1,000 repositories—and task‑oriented skills tailored to specific research tasks. When equipped with these skills, the agent achieves significant performance gains across multiple benchmarks, outperforming a skill‑free version by 134.3% on MLE‑bench, 34.4% on PaperBench, 9.2% on FrontierCS, and 14.0% on PassNet.

arXiv Computation and Language
Sep 11

INDRA: A New AI Tool for Exploring Tobacco, Fossil Fuel, and Chemical Industry Archives

INDRA is a research platform that integrates multiple archival collections—such as UCSF’s Industry Documents Library, Columbia and CUNY’s ToxicDocs, and Stanford’s SRITA—into a single, LLM‑readable corpus. It employs three safeguards: a closed evidentiary sandbox, real‑time provenance tagging, and a deterministic system‑level protocol to ensure that model outputs are clearly distinguished from archival evidence and from the model’s own inferences. The platform enables large‑language‑model‑powered investigations across these archives while keeping the conditions of knowledge production transparent and auditable.

By Daniel Akselrad, Robert N. Proctor