arXiv AI By Yuchen Xia, Michael Weyrich, Nasser Jazdi, Johannes St\"umpfle, Johannes Sigel, Akshay Narla, Gavin K. Reynolds, Anna Jawor-Baczynska, Pol Llopart

LLM Agents Perform Controlled Experiments Using Simulation Models

Read the original on arXiv AI →

The paper introduces a multi‑agent framework that lets large language models (LLMs) perform controlled experiments using scientific simulation models, specifically for pharmaceutical process design. Given a user query and baseline configuration, the system builds a structured task, designs and runs comparative simulations, interprets outcomes, and generates evidence‑based recommendations for optimizing process parameters. By integrating high‑fidelity simulations with LLMs, the approach yields more specific, actionable outputs and improves user‑rated correctness and helpfulness compared to language‑only reasoning.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 22

Agents in the Wild: Where Research Meets Deployment

arXiv:2607. 19336v1 Announce Type: new Abstract: Agentic systems large language model (LLM) based architectures capable of reasoning, planning, acting, and coordinating with tools and other agents are rapidly transitioning from research prototypes to production scale deployments across domains such as software engineering, scientific discovery, and finance.

By Grace Hui Yang, Pranav N. Venkit, Hooman Sedghamiz, Enrico Santus, Victor Dibia, Ioana Baldini
arXiv AI
6d ago

From Atomic to Agentic: Towards Interpretable Evaluation of LLMs' Agentic Mathematical Capabilities

The paper introduces a new benchmark that evaluates large language models (LLMs) on their agentic mathematical reasoning rather than just final answers. It aligns problem‑solving behaviors with a taxonomy of reusable mathematical atomic capabilities and includes planning, action, and feedback tasks in both textual and multimodal settings. Experiments show that models with similar end‑to‑end accuracy can have very different agentic profiles, highlighting the importance of process‑level evaluation.

By Jiayi Kuang, Yinghui Li, Yunze Song, Keyu Chen, Zhifeng Shen, Yangning Li, Yidong Wang, Di Yin, Ruizhi Qiao, Xing Sun, Kai Jin, Ying Shen, Liang Lin, Philip S. Yu