arXiv AI

MemMux: Runtime Verification and Honest Resource Attribution for Fleets of Parallel Coding Agents

MemMux is a local runtime designed to provide runtime verification and honest resource attribution for fleets of parallel coding agents. It emits observable signals that track per‑agent memory usage, ensure complete reclamation of terminated agents, detect escaped child processes, and keep the system from exceeding a bounded memory footprint. In benchmarks against tmux and a raw‑process baseline, MemMux keeps a fleet under a 7.5 GiB budget with zero swap, while ungoverned tools exceed the budget and spill into swap, and it achieves 100 % attribution with low overhead.

Hugging Face Trending Papers
Jun 2

Agent libOS: A Library-OS-Inspired Runtime for Long-Running, Capability-Controlled LLM Agents

Large language model (LLM) agents are evolving from request-response assistants into long-running software actors: they maintain state across model calls, fork subtasks, wait for external events, request human authority, generate tools, and perform side effects that must be resumed and audited. This paper presents Agent libOS, a library-OS-inspired runtime substrate for LLM agents.

arXiv AI
Aug 19

SkillEffect: Checked Lowering for Memory-Bounded Agent Tools

SkillEffect is a checked‑lowering runtime that ensures agent tool calls stay within memory limits by verifying each proposed program against an immutable input before execution. It uses audited relation plugins to provide source recognition, bounded intermediate representation construction, and postconditions, while a shared runtime handles selection, bounded VM execution, and atomic capacity leasing. Experiments across six operator families show that bounded access significantly reduces peak memory usage and improves completion rates under fixed memory caps.

By Yinuo Wang, Yiyu Shi
arXiv AI
Aug 26

Rebuild Dossier: Mechanically-Enforced Specs for Agentic App Rebuilds, and What Model-Tier Failures Reveal

The paper introduces rebuild‑dossier, an open‑source tool that locks an application’s real interface before code is written and enforces one‑test‑at‑a‑time building through automated checks. In experiments, a compliant agent failed a held‑back test while a rule‑breaking agent passed, showing that a passing test suite can be gamed. The study also demonstrates that the automated check mechanism, rather than interface‑locking alone, is crucial for reliable rebuilds, and that multi‑level verification catches errors that single‑level checks miss.

By Parker Fawcett
arXiv AI
3d ago

Inherit-MAS: Test-Time Evolution of Multi-Agent Systems through Workflow and Execution Inheritance

Inherit-MAS introduces explicit inheritance at both workflow and execution levels for multi‑agent systems built from large language models. A meta‑model generates a workflow of specialized agents, while a judge evaluates each candidate and diagnoses deficiencies. During refinement, workflow inheritance discards unhelpful nodes and applies validated edits, and execution inheritance reuses stored results only when the request and context match, reducing redundant computation.

By Songtao Wei, Yi Li, Zhichun Guo, Bingzhe Li
arXiv AI
1d ago

Verifying Coordination in Parallel Coding Agents: NP-Bench and a Scheduling Planner

The paper introduces NP‑Bench, a benchmark and a proactive scheduling planner that coordinates parallel large‑language‑model coding agents. By partitioning work scopes and ordering merges ahead of time, the planner improves clean‑integration rates from 1/9 to 9/9 and eliminates merge conflicts, outperforming both no‑coordination and reactive‑detection baselines. It also demonstrates that cross‑session memory can eliminate repeated mistakes and that routing facts to agents does not improve long‑context accuracy at scale.

By Sumanyu Muku