What We Learned by Reproducing 2,200 papers from ICML
Read the original on Hugging Face Blog →The Flow has not summarised this story yet — read it at Hugging Face Blog.
The Flow has not summarised this story yet — read it at Hugging Face Blog.
A topic-organized collection of 200+ LLM research papers from 2025
A curated roundup of notable LLM research papers that came out this year
Modelpedia is an automated, LLM-assisted framework that extracts and organizes findings about AI models from published papers into a searchable public catalog. It links each finding to the relevant model, dataset, method, and concept, and has already extracted over a thousand findings from ICLR 2024 and 2025 papers. The authors invite the community to explore, contribute to, and build on this open catalog, positioning model findings as a shared foundation for the meta‑science of AI.
arXiv:2604.21965v2 Announce Type: replace Abstract: Recent work has used LLM agents to reproduce empirical social science results with access to both the data and code. We broaden this scope by askin...
arXiv:2607. 28618v1 Announce Type: cross Abstract: Chemistry literature synthesis often requires assembling specific findings scattered across many publications, yet existing literature-search systems primarily return ranked document lists.
The article introduces AgentActionBench, a benchmark designed to evaluate agent-based experiment reproduction across machine learning and AI4Science papers. It employs an MCP-based Action Recorder to capture agents’ behavior during reproduction and assesses the resulting traces against paper-specific rubrics. The benchmark includes 150 papers, with a human-annotated subset and model-assisted augmentation expanding it to over 10,000 rubric items, revealing that current systems face execution bottlenecks but that model-generated rubrics correlate strongly with human judgments.