arXiv Computation and Language By Jiasheng Shi, Tianhan Zhang

What Do CAE Simulation Agents Really Need Beyond a Generic Harness?

Read the original on arXiv Computation and Language →

The paper investigates whether computer‑aided engineering (CAE) simulation agents require specialized features beyond a generic LLM harness. Experiments show that a single‑agent harness with multi‑turn reasoning, tool use, and execution feedback can match or outperform multi‑agent specialized systems, with domain knowledge (solver tutorials) providing the most significant performance boost. The study highlights that modern harnesses already supply many capabilities previously thought necessary for CAE agents.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Jun 15

HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry

arXiv:2606. 14249v1 Announce Type: new Abstract: AI agent performance depends critically on the runtime harness, comprising the prompts, tools, memory, and control flow that mediate how a model observes, reasons, and acts.

By Tingyang Chen, Shuo Lu, Kang Zhao, Weicheng Meng, Hanlin Teng, Tianhao Li, Chao Li, Xule Liu, Jian Liang, Zhizhong Zhang, Yuan Xie, Heng Qu, Kun Shao, Jian Luan
arXiv AI
Aug 14

Foam-Agent: A Large Language Model-Based Multi-Agent Framework for Automating Computational Fluid Dynamics Workflows

arXiv:2505. 04997v3 Announce Type: replace Abstract: Computational fluid dynamics (CFD) has been the main workhorse of computational physics, yet its steep learning curve and fragmented, multi-stage workflow create significant barriers to entry.

By Ling Yue, Nithin Somasekharan, Tingwen Zhang, Yadi Cao, Zhangze Chen, Shimin Di, Shaowu Pan
arXiv AI
Aug 26

Beyond Executable Models: The Pufibara Agent Harness and the Modelica Agent Workflow Benchmark for Physical System Modeling

The paper introduces Pufibara, an agent harness designed to maintain engineering state and evidence across revisions in Modelica-based physical system modeling. It also presents a 232-task Modelica Agent Workflow Benchmark covering model repair, generation, and tuning, evaluated by an external benchmark-owned evaluator. Experiments show Pufibara outperforms Claude Code in task success and resource efficiency across two LLM backends.

By Zizhe Wang
arXiv AI
Sep 25

Grow the Harness, Not the Context: From Strategy-Free Scaffolds to Reusable Specialist Agents

The paper introduces Growing Harness, a training method that transforms recurring control logic in large language model agents into reusable executable code, reducing reliance on the model for task-specific decisions. By using strategy-free scaffolds, failure-guided code repair, and success-first gating, the approach learns a shared harness that improves performance across multiple benchmarks and model sizes. Experiments on BrowseComp-Plus and WebArena-Verified show significant gains in success rates and substantial reductions in LLM calls and inference cost compared to traditional tool‑calling agents.

By Laizhen Li, Jiarui Li, Juanjuan Zhao, Kejiang Ye, Ye Li, Cheng-zhong Xu, Xitong Gao
arXiv Machine Learning
Aug 27

JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution

JIT‑Agent is a model that automatically generates task‑adaptive agent harnesses for any off‑the‑shelf LLM, replacing manual, task‑specific harness design. It learns to compose, repair, and evolve harnesses using a fixed four‑module protocol, and its use boosts performance on benchmarks such as DeepSearchQA and OdysseyBench, outperforming several mature agent runtimes. The approach demonstrates that harness intelligence can be trained, transferred, and compounded independently of model scaling.

By Guibin Zhang, Leo Lu, Fangzhou Xie, Kang Zhu, Junhao Wang, Zhifei Xie, Zhaochen Yu, Zihang Liu, Zhongxiang Sun, Qiankun Li, Yue Liao, Heng Chang, Xiaobin Hu, Qibing Ren, Wangchunshu Zhou, Shuicheng Yan
arXiv AI
Sep 25

MOOSEnger: A Simulation-Aware AI Agent Framework for the MOOSE Ecosystem

MOOSEnger is a simulation‑aware AI agent framework designed for the MOOSE ecosystem, integrating an interchangeable reasoning model with domain knowledge, revised simulation artifacts, MOOSE‑specific validation, and executable solver feedback. Its generate‑check‑repair‑run workflow uses MOOSE knowledge retrieval, HIT‑aware parsing, syntax metadata, diagnostics, and revision‑controlled authoring to bind evidence to each input revision and guide bounded repair before acceptance. Across 200 prompts, MOOSEnger raises executable success from 5% to 89.5% with GPT‑5.2 and from 0% to 76.5% with Gemma 4 31B, and a ten‑case benchmark shows all generated inputs meet semantic alignment, with eight also meeting numerical‑accuracy criteria.

By Mengnan Li, Jason Miller, Zaid Abulawi, Zachary Prince, Matt Kohl, Jack M. Cavaluzzi, Guillaume Giudicelli, Casey T. Icenhour, Alexander Lindsay, Cody Permann