Read the Paper, Write the Code: Agentic Reproduction of Social-Science Results
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2609.22111v1 Announce Type: new Abstract: Large language model agents are increasingly capable of conducting research autonomously, producing research documents alongside the code and experimen...
The paper introduces AutoMat, a benchmark designed to test large language model (LLM) coding agents on their ability to reproduce claims from computational materials science. AutoMat presents three challenges: reconstructing underspecified procedures, navigating specialized toolchains, and assessing whether the evidence supports a claim. Experiments show that current LLM agents achieve low success rates, with the best setting reaching only 53%, and failures stem mainly from incomplete procedures, methodological deviations, and execution fragility.
arXiv:2606. 18237v1 Announce Type: cross Abstract: Reproducing research results from papers and released code is central to scientific progress.
arXiv:2602. 11354v3 Announce Type: replace Abstract: The literature has witnessed an emerging interest in AI agents for automated assessment of scientific papers.
arXiv:2607. 02134v1 Announce Type: new Abstract: Scientific machine learning papers typically make computational claims, e.
arXiv:2609.01526v1 Announce Type: new Abstract: Scientific agents must learn not only how to reason, but also what to believe. However, existing LLM agents typically express scientific hypotheses in...