arXiv AI

One Step at a Time: Trading LLM Autonomy for Process Predictability

arXiv AI
2d ago

VERSE: Verified Self-Evolving Optimizer for Agent Harnesses

The paper introduces VERSE, a Verified Self‑Evolving optimizer that enhances LLM agent harnesses by allowing the optimizer to test edits, replay failures, and perturb steps while tracking fixes and regressions. VERSE builds its own tools for failure analysis, verification, training audits, and workflow control, and uses this feedback to revise the harness’s prompts, skills, tools, hooks, and notes without changing model weights. In experiments across five executors and multiple languages, VERSE improves all evaluated harness optimizers, achieving higher accuracy on held‑out and out‑of‑distribution tasks compared to the strongest baselines.

By Zekai Wang, Yingqiang Ge, Zekun Wang, Hai Wang, Yuhui Xu, Joshua Frandsen, Shancong Fu, Ashia C. Wilson, Chandan K. Reddy
arXiv AI
Sep 25

Policy as Code: A Coroutine-Bridge Harness for Fast-Reasoning Reliability on CAR-bench

The paper introduces a coroutine-bridge harness that lets a language model emit a Python program to manage tool calls in the CAR-bench evaluation. By decoupling model invocations from tool round-trips, the approach reduces model calls to a median of two per task while maintaining seven agent turns, achieving a median latency of 1.8 s on a Cerebras gpt‑oss‑120b. The harness achieved 60.0 % Pass³ on the official hidden evaluation, outperforming the baseline by 4.5× and matching frontier-model agents on GPT‑5.5, all while keeping the prompt largely cached and minimizing input compute.

By Ivan Matveev
arXiv AI
Sep 7

FinalityBench: An Effect-Level Benchmark for Agent Decisions Under Delayed and Conflicting Financial Finality

FinalityBench is an executable benchmark that tests how agents decide on shipping, re‑capturing, refunding, or waiting when a merchant’s payment processor, ledger, ERP, and bank feed receive delayed, duplicated, dropped, or reordered messages, causing contradictory beliefs about an order. The benchmark uses a hidden canonical event log and faulted delivery streams to generate system views, scoring each episode by the merchant’s terminal economic position relative to a privileged reference. It contains 321 tasks, including 45 twin pairs where all four views are identical yet the correct disposition differs, and evaluates nine programmatic policies, revealing that a ship‑on‑first‑sign policy performs best by accuracy but worst by paired loss, while a runtime‑gated irreversible‑action policy achieves 85.4% accuracy without losing money.

By Abhishek Sharma
arXiv AI
Sep 2

Long-Horizon State Tracking in LLMs: Executing MD5 through a Deep Sequence of Dependent Tool Calls

The paper introduces a benchmark for evaluating large language models (LLMs) on long‑horizon state tracking by having them compute the MD5 hash through 196 dependent tool calls across 64 rounds, carrying four 32‑bit words in context. It shows that a mixture‑of‑experts LLM can maintain the full state and produce correct digests in most runs, even when all primitive tools are replaced by another LLM. The study isolates state‑tracking difficulty from instruction interpretation and identifies key factors—contextual reasoning and worker voting—that enable success.

By Dheeraj Mohandas Pai, Lu Xian