WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
AI research agents need reliable knowledge of how their experiments change outcomes. We introduce WhatWorkedBench to measure experimental understanding, the accuracy of predictions about component cha...
arXiv:2608.31076v1 Announce Type: cross Abstract: Autonomous scientific research agents are increasingly applied to end-to-end scientific workflows, including literature review, data analysis, experi...
arXiv:2606. 26158v1 Announce Type: new Abstract: When a benchmark's accuracy saturates, it is often retired and replaced with a more challenging version.
OpenDiscoveryTrace is a public dataset of 558 complete AI scientific agent trajectories that records the reasoning process—thoughts, tool calls, observations, errors, revision triggers, and confidence—across 124 scientific tasks in drug discovery, materials science, genomics, and literature analysis. The dataset includes seven models (three frontier models and four open‑weight models) and 60 live‑retrieval variants, providing a balanced view of performance and error patterns. Pilot analysis shows that process traces reveal behavioral differences invisible to output‑only evaluation, such as differing error rates and types among frontier models.
arXiv:2608. 15089v1 Announce Type: new Abstract: Long-horizon agents can fail even when their underlying models can solve the constituent steps.
The paper investigates how altering the harness—specifically the way a coding agent manages context and tool outputs—affects performance when the underlying model and task remain unchanged. Two harness configurations were compared on three coding benchmarks: a control that preserves the full conversation in order, and a treatment that mechanically shortens older tool results to keep the context tight. Across all benchmarks, the treatment increased the mean per‑task fail‑to‑pass fraction and, in some cases, the number of complete solutions, demonstrating that the harness itself can significantly influence a frozen model’s effectiveness.