Hugging Face Trending Papers

Can LLM design high-quality experiments? A Comprehensive and Systematic Benchmark on Autonomous Experimental Design

AI for Research (AI4Research) leverages AI to automate and improve scientific workflows. While experimental design is a critical stage of the research process, prior work has focused primarily on code implementation and execution, overlooking the importance of this stage, and no benchmark exists to evaluate AI's ability to conduct systematic experiment design.

arXiv AI
Sep 16

Multi-Agent Collaboration for Automated Design Exploration on High Performance Computing Systems

The paper introduces MADA, a Large Language Model–powered multi‑agent framework that coordinates specialized agents—Job Management, Geometry, and Inverse Design—to automate complex design workflows on high‑performance computing systems. In the context of Richtmyer–Meshkov Instability suppression for Inertial Confinement Fusion, MADA iteratively refines designs by launching ensemble simulations, generating meshes, and proposing new designs based on simulation outcomes, achieving improved suppression with minimal manual effort. The framework demonstrates how coordinated reasoning, simulation, and specialized tools can be scaled for rapid, automated design exploration.

By Harshitha Menon, Charles F. Jekel, Kevin Korner, M. Giselle Fernandez-Godino, Brian Gunnarson, Nathan K. Brown, Michael Stees, Walter Nissen, Meir H. Shachar, Dane M. Sterbentz, William J. Schill, Yue Hao, Robert Rieben, William Quadros, Steve Owen, Scott Mitchell, Ismael D. Boureima, Jonathan L. Belof
arXiv AI
Sep 7

La Agente \'Optima: Towards Agentic Self-Driving Laboratories

La Agente ’Optima is an agentic framework that builds and manages Bayesian optimization campaigns for self‑driving laboratories, separating large language model reasoning from campaign execution. It maintains a persistent optimization state, allowing consistent repetitive loops and auditable decisions, and only returns control to the agent when interpretation or revision is needed. In tests on digital discovery tasks and physical platforms, it corrected measurement failures, improved yields, and recommended formulation changes, outperforming human‑directed campaigns in cost and material usage.

By Marcel M\"uller, Jiaru Bai, Willi Gottstein, Abhijoy Mandal, Mohammad Nazeri, Elia Savino, Yanlin Fang, Sujoy Das, Sergio Pablo Garc\'ia Carrillo, Yeonghun Kang, Juan B. P\'erez-S\'anchez, Simone Pilon, Martin Fitzner, Timothy No\"el, Frank Gu, Varinia Bernales, Al\'an Aspuru-Guzik
arXiv Machine Learning
Jun 25

ASAP: Agent-System Co-Design for Wall-Clock-Centered Auto HPO Research for ML Experiments

arXiv:2606. 25207v1 Announce Type: new Abstract: Hyperparameter Optimization (HPO) is essential for maximizing machine learning model performance, and its core challenge is sample efficiency: finding strong configurations within a limited budget.

By Taicheng Guo, Haomin Zhuang, Kehan Guo, Yujun Zhou, Nitesh V. Chawla, Olaf Wiest, Xiangliang Zhang
arXiv AI
2d ago

CompMat-Bench: Benchmarking AI Agents for Computational Materials Science

CompMat-Bench is a new benchmark comprising 94 tasks drawn from recent computational materials science studies, designed to evaluate AI agents on realistic research steps without requiring costly simulations during testing. The benchmark pre‑reproduces inputs and outputs to provide ground truth, allowing agents to be graded with fixed rules rather than an LLM judge. It supports both single tasks and multi‑step workflows, with varying levels of methodological guidance, and shows that while agents can achieve high pass rates on individual tasks, performance drops in longer workflows or with reduced guidance, often due to scientific rather than software errors.

By Chenmu Zhang, Levi Felix, Jun-Jie Zhang, Xingfu Li, Xuelian Jiang, Tao Jiang, Subhendu Mishra, Xixi Qin, Boris Yakobson
arXiv Machine Learning
Sep 22

Strategy Accumulation and Guided Execution for Automated LLM Fine-Tuning

The paper introduces Strategy Accumulation and Guided Execution (SAGE), a two-stage framework that makes automated fine-tuning of large language models cumulative. In the first stage, a multi-agent pipeline uses Monte Carlo Tree Search to explore training strategies while a Distillation Agent records task-specific insights and cross-task confidence scores into a structured repository. In the second stage, SAGE retrieves relevant experience from this repository to guide training on new tasks, achieving a 12.4‑percentage‑point improvement over a baseline pipeline without accumulated experience on nine unseen tasks.

By Haoran Zhao, Wei Du, Dingwen Yang, Jixuan Huang, Junlin Shang, Lingyong Fang, Ya Guo, Tao Gui, Qi Zhang, Xuanjing Huang
arXiv AI
Aug 28

AgentFold: Closed-Loop Agentic Search for Protein Folding Model Design

AgentFold is a multi‑agent framework that treats protein‑folding model design as a closed‑loop search over executable code variants. Starting from the ESMFold codebase, the agents generate hypotheses, modify and debug code, evaluate model variants, and store both successes and failures in structured memory, guided by an MCTS‑style policy that allocates GPU resources. In an engineering‑scale experiment, AgentFold explored about 80 variants using 5,000 GPU‑hours and 170 million LLM tokens, improving the best lDDT score by 7.5% over independent Codex proposals and outperforming a random‑search baseline, while also uncovering empirical design patterns such as the benefits of early, soft, learnable priors.

By Mingquan Liu, Jiangyu Chen, Hanqun Cao, Xujun Zhang, Pengsen Ma, Xiangru Tang, Shuting Jin, Zhuo Yang, Tianfan Fu, Fang Wu, Xiangxiang Zeng
arXiv Machine Learning
Jul 17

Idea2Plan: Exploring AI-Powered Research Planning

arXiv:2510. 24891v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have demonstrated significant potential to accelerate scientific discovery as valuable tools for analyzing data, generating hypotheses, and supporting innovative approaches in various scientific fields.

By Jin Huang, Silviu Cucerzan, Sujay Kumar Jauhar, Ryen W. White