arXiv AI By Yongchao Ye, Xinyu He, Dutliff Boshoff, Way Kuo, Lishuai Li

little m: An AI Agent for Industrial Process Optimization

Read the original on arXiv AI →

The paper introduces little m, an AI agent that helps formulate industrial process control models by combining a domain-specific knowledge repository with LLM-driven interaction. It tackles the challenge of converting messy real-world specifications, including natural language and spatial diagrams, into rigorous mathematical optimization models. The authors also present IPC-Bench, a multimodal dataset of 50 canonical scenarios, and show through automated and human evaluations that little m outperforms state‑of‑the‑art LLMs in generating semantically correct models.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Jun 9

GPT-Micro: A large language paradigm for accelerated, inexpensive, and thermodynamics-consistent discovery of constitutive models in manufacturing

arXiv:2606. 08238v1 Announce Type: new Abstract: Constitutive modeling of the relationship between process-imposed material states and fundamental material properties is critical to control of material microstructure in manufacturing processes.

By Soumik Dutta, Kiarash Naghavi Khanghah, Sania Shree, Logan McNeil, Thomas Feldhausen, Hongyi Xu, Rajiv Malhotra
arXiv AI
Aug 26

LLM Agents Perform Controlled Experiments Using Simulation Models

The paper introduces a multi‑agent framework that lets large language models (LLMs) perform controlled experiments using scientific simulation models, specifically for pharmaceutical process design. Given a user query and baseline configuration, the system builds a structured task, designs and runs comparative simulations, interprets outcomes, and generates evidence‑based recommendations for optimizing process parameters. By integrating high‑fidelity simulations with LLMs, the approach yields more specific, actionable outputs and improves user‑rated correctness and helpfulness compared to language‑only reasoning.

By Yuchen Xia, Michael Weyrich, Nasser Jazdi, Johannes St\"umpfle, Johannes Sigel, Akshay Narla, Gavin K. Reynolds, Anna Jawor-Baczynska, Pol Llopart
arXiv AI
Aug 28

AI Control Scientist: LLM-driven Agentic System for Automated Control Design

AI Control Scientist (AICS) is a large language model–driven agent that automatically generates optimized controllers from language design requirements. It comprises a Task Modeling Agent that translates user needs into engineering constraints, a Controller Design Agent that produces candidate controller structures and code, and a Parameter Tuning Agent that refines parameters to meet closed‑loop performance criteria. Experiments show AICS outperforms existing automated baselines in design success rate and optimization efficiency, enabling the creation of multiple representative control systems.

By Haiteng Wang, Weihao Li, Jing Zhang, Lei Ren
arXiv Computation and Language
Aug 25

MIRROR: A Multi-Agent Framework with Iterative Adaptive Revision and Hierarchical Retrieval for Optimization Modeling in Operations Research

arXiv:2602.03318v4 Announce Type: replace Abstract: Operations Research (OR) relies on expert-driven modeling--a slow and fragile process ill-suited to novel scenarios. While large language models (L...

By Yifan Shi, Jiayi Wang, Minyi Wu, Ye Fan, Jialong Shi, Jianyong Sun
arXiv AI
Aug 28

From Atomic to Agentic: Towards Interpretable Evaluation of LLMs' Agentic Mathematical Capabilities

The paper introduces a new benchmark that evaluates large language models (LLMs) on their agentic mathematical reasoning rather than just final answers. It aligns problem‑solving behaviors with a taxonomy of reusable mathematical atomic capabilities and includes planning, action, and feedback tasks in both textual and multimodal settings. Experiments show that models with similar end‑to‑end accuracy can have very different agentic profiles, highlighting the importance of process‑level evaluation.

By Jiayi Kuang, Yinghui Li, Yunze Song, Keyu Chen, Zhifeng Shen, Yangning Li, Yidong Wang, Di Yin, Ruizhi Qiao, Xing Sun, Kai Jin, Ying Shen, Liang Lin, Philip S. Yu