arXiv AI By Kyoungmin Kim, Anastasia Ailamaki

Confining Nondeterminism: AI-Driven Research Systems as DBMSs for Reliable, Non-Wasteful, Transparent, and Collaborative Research [Vision]

Read the original on arXiv AI →

arXiv:2607. 10508v1 Announce Type: cross Abstract: LLM agents that conduct research (proposing ideas, writing and running code, analyzing results) can already carry a study from research question to figures, yet cannot be fully trusted.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 11

Introducing Consort: A Spec-First Agent Framework for Enforced, Test-Driven Development on Live Database Branches

Consort is a spec‑first, test‑driven agent framework that enforces engineering discipline through immutable controls, a deterministic orchestrator, and human‑approved gates. It operates on live database branches, guiding role agents through a spec‑first design lane and a test‑driven build lane. The framework claims that enforcing tests and gates in code keeps agent‑written code honest and verifiable, while specialized roles make it maintainable.

By Kevin Hartman
arXiv AI
4d ago

StateTape: Action-Conditioned Evidence Lifecycle Modeling for Long-Horizon Coding Agents

StateTape introduces a new framework for long‑horizon coding agents that rewrites the agent’s context as the code repository changes, rather than letting the context grow with every observation. It models the repository as a symbol‑level code graph, using a tape to mark symbols altered by each write and a manager model to resolve stale records. The authors provide theoretical analysis, a new benchmark called TraceBench, and empirical results showing higher resolve rates across six agents and three edit‑heavy benchmarks with minimal computational overhead.

By Ziyang Yu, Liang Zhao, Bowen Zhu, Hasibul Haque