arXiv AI

GitLake: Git-for-data for the agentic lakehouse

arXiv:2607. 08319v1 Announce Type: cross Abstract: We present GitLake, a Git-for-data design for an agent-first lakehouse.

arXiv AI
Sep 3

Git4Data: Database-Native Version Control for AI Agents

Git4Data introduces a database-native version‑control layer that treats a database as a repository and each table as a versioned object, exposing Git‑style operations—snapshot/tag, branch, diff, and merge—through SQL extensions. Implemented in MatrixOne, it leverages immutable object storage and MVCC so that operation costs depend on the size of the change rather than the entire dataset. In agentic branching workloads, Git4Data outperforms DoltDB by up to an order of magnitude, demonstrating efficient versioning for AI agents.

By Hongshen Gou, Zuyu Zhang, Yuze Sun, Peng Xu, Feng Tian, Long Wang, Jianguo Wang
arXiv AI
Jun 2

"Skill issues'': data-centric optimization of lakehouse agents

arXiv:2606. 01185v1 Announce Type: new Abstract: Coding agents are becoming users of data infrastructure, but their success depends not only on model quality: it also depends on the skills and environment files that teach agents how to use a system.

By Nicole Rose Schneider, Davide Ghilardi, Giacomo Piccinini, Jacopo Tagliabue
arXiv AI
Jul 17

"Skill Issues'': Data-Centric Optimization of Lakehouse Agents

arXiv:2606. 01185v2 Announce Type: replace Abstract: Coding agents are becoming users of data infrastructure, but their success depends not only on model quality: it also depends on the skills and environment files that teach agents how to use a system.

By Nicole Rose Schneider, Davide Ghilardi, Giacomo Piccinini, Jacopo Tagliabue
arXiv AI
Sep 25

AgileLog: A Forkable Shared Log for Agents on Data Streams

AgileLog introduces a forkable shared log designed to support AI agents that interact with streaming data. The new abstraction provides forking primitives that allow agents to operate without causing performance interference or unsafe writes. Bolt is a system that implements AgileLog, employing techniques to keep forks inexpensive while ensuring logical and performance isolation.

By Shreesha G. Bhat, Tony Hong, Michael Noguera, Aishwarya Ganesan, Ramnatthan Alagappan
arXiv AI
Jun 4

Archi: Agentic Operations at the CMS Experiment

arXiv:2606. 04755v1 Announce Type: cross Abstract: We present Archi, an open-source, end-to-end framework for scientific collaborations that combines the systematic ingestion and organization of heterogeneous data sources with the deployment of configurable, private, and extensible agents that retrieve and reason over them.

By Pietro Lugato, Luca Lavezzo, Jason Mohoney, Hasan Ozturk, Muhammad Hassan Ahmed, Juan Pablo Salas, Viphava Ohm, Krittin Phornsiricharoenphant, Gabriele Benelli, Mariarosaria D'Alfonso, Manasvita Joshi, Warren Nam, Aron Soha, Samantha Sunnarborg, Austin Swinney, Jack Tucker, Dmytro Kovalskyi, Tim Kraska, Christoph Paus
arXiv AI
Sep 3

SpecMine: A Large-Scale Corpus of Spec-Driven Development Artifacts

SpecMine is a large-scale corpus that documents Spec-Driven Development (SDD) artifacts in public GitHub repositories. It includes a broad census of 470,795 spec files from 73,030 repositories linked to 17 tools, a focused census of 98,574 Kiro layout files from 12,910 repositories, and a sweep of 5,992 pull requests across 581 repositories that modify specs. The dataset provides enriched metadata, full commit histories, parsed document structures, and over 2.4 million typed references connecting specs to code, sibling documents, PRs, branches, and issues.

By Shyam Agarwal, Anmol Singhal, Travis Breaux, Bogdan Vasilescu