WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
arXiv:2609.27490v1 Announce Type: new Abstract: AI research agents need reliable knowledge of how their experiments change outcomes. We introduce WhatWorkedBench to measure experimental understanding...
arXiv:2608.31076v1 Announce Type: cross Abstract: Autonomous scientific research agents are increasingly applied to end-to-end scientific workflows, including literature review, data analysis, experi...
Autonomous scientific research agents are increasingly applied to end-to-end scientific workflows, including literature review, data analysis, experimentation, and report generation. However, open-end...
The paper investigates how altering the harness—specifically the way a coding agent manages context and tool outputs—affects performance when the underlying model and task remain unchanged. Two harness configurations were compared on three coding benchmarks: a control that preserves the full conversation in order, and a treatment that mechanically shortens older tool results to keep the context tight. Across all benchmarks, the treatment increased the mean per‑task fail‑to‑pass fraction and, in some cases, the number of complete solutions, demonstrating that the harness itself can significantly influence a frozen model’s effectiveness.
OpenDiscoveryTrace is a public dataset of 558 complete AI scientific agent trajectories that records the reasoning process—thoughts, tool calls, observations, errors, revision triggers, and confidence—across 124 scientific tasks in drug discovery, materials science, genomics, and literature analysis. The dataset includes seven models (three frontier models and four open‑weight models) and 60 live‑retrieval variants, providing a balanced view of performance and error patterns. Pilot analysis shows that process traces reveal behavioral differences invisible to output‑only evaluation, such as differing error rates and types among frontier models.
arXiv:2602. 11354v3 Announce Type: replace Abstract: The literature has witnessed an emerging interest in AI agents for automated assessment of scientific papers.