Hugging Face Blog

What We Learned by Reproducing 2,200 papers from ICML

arXiv Machine Learning
Sep 2

Modelpedia: A Catalog of Model Findings for the Meta-Science of AI

Modelpedia is an automated, LLM-assisted framework that extracts and organizes findings about AI models from published papers into a searchable public catalog. It links each finding to the relevant model, dataset, method, and concept, and has already extracted over a thousand findings from ICLR 2024 and 2025 papers. The authors invite the community to explore, contribute to, and build on this open catalog, positioning model findings as a shared foundation for the meta‑science of AI.

By Franciszek Bernat (Centre for Credible AI, Warsaw University of Technology), Dawid P{\l}udowski (Centre for Credible AI, Warsaw University of Technology), Micha{\l} Jan W{\l}odarczyk (Centre for Credible AI, Warsaw University of Technology), Luca Longo (University College Cork), Jianlong Zhou (University of Technology Sydney), Andreas Holzinger (Human-Centered AI Lab), Riccardo Guidotti (University of Pisa, ISTI-CNR), Wojciech Samek (Technical University of Berlin, Berlin Institute for the Foundations of Learning and Data), Przemys{\l}aw Biecek (Centre for Credible AI, University of Warsaw)
arXiv Computation and Language
Sep 11

Overview of the NLPCC 2026 Shared Task 11: Agent-Based Experiment Reproduction from Scientific Papers

The article introduces AgentActionBench, a benchmark designed to evaluate agent-based experiment reproduction across machine learning and AI4Science papers. It employs an MCP-based Action Recorder to capture agents’ behavior during reproduction and assesses the resulting traces against paper-specific rubrics. The benchmark includes 150 papers, with a human-annotated subset and model-assisted augmentation expanding it to over 10,000 rubric items, revealing that current systems face execution bottlenecks but that model-generated rubrics correlate strongly with human judgments.

By Hanhua Hong, Yizhi Li, Luu Gia Huy, Jian Yang, Ming Zhou, Chenghua Lin
arXiv AI
Aug 26

ReproAgent: Contract-Guided Paper-to-Code Reproduction

ReproAgent is a four‑stage pipeline—Prepare, Plan, Generate, Repair—that uses a persistent implementation contract to guide scientific AI agents in converting research papers into executable code repositories. The system employs two channels: an implementation‑requirement channel that translates paper snippets into code obligations, and a reference‑evidence channel that pulls content and structure from related repositories. Evaluated on PaperBench Code‑Dev, ReproAgent achieves the highest mean score among same‑backbone scaffolds for both Claude‑Sonnet‑4.5 and Gemini‑3‑Flash, with ablation studies confirming the contribution of both channels.

By Xue Hu, Zewei Pan, Zhongyuan Wang, Zhou Liu, Zeli Su, Wentao Zhang
arXiv Computation and Language
5d ago

Drift Inspector: Exploring and Measuring Scientific Drift with Atomic Contribution Claims

Drift Inspector is an open‑source system that extracts Atomic Contribution Claims (ACCs) from scientific abstracts using an LLM, then clusters these claims over time to map how a research field evolves. Applied to six years of EMNLP, the tool reveals a shift from classic NLP tasks toward LLM‑era capabilities such as reasoning and multimodality—trends that keyword or whole‑abstract counts miss. The pipeline has also processed the entire ACL Anthology, yielding 346,000 claims from 80,000 abstracts across 423 venues, with human‑validated extraction and clustering aligned to an external taxonomy.

By Vsevolod Karimov, Stepan Ostarkov, Anastasia Poroshina, Anatoly Frolov, Alexander Panchenko