arXiv AI

AURORA: A Natural Language-Driven Agentic Framework for Understanding, Reasoning, and Orchestrating Reliable Air-Ground Co-Simulation

AURORA is a natural‑language‑driven framework that treats air‑ground scenario generation as a compilation process with verification. It introduces the Air‑Ground Scenario Graph (AGSG), a typed intermediate representation linking agents, missions, events, communication, and success conditions, enabling joint grounding, temporal planning, pre‑execution checks, runtime verification, failure localization, and bounded repair. The authors also present AURORA‑Bench to evaluate not only execution but faithful realization of requested interactions, showing that structured execution and runtime verification improve reliability and that explicit intermediate representations facilitate verifiable and repairable co‑simulation.

arXiv AI
Sep 25

MOOSEnger: A Simulation-Aware AI Agent Framework for the MOOSE Ecosystem

MOOSEnger is a simulation‑aware AI agent framework designed for the MOOSE ecosystem, integrating an interchangeable reasoning model with domain knowledge, revised simulation artifacts, MOOSE‑specific validation, and executable solver feedback. Its generate‑check‑repair‑run workflow uses MOOSE knowledge retrieval, HIT‑aware parsing, syntax metadata, diagnostics, and revision‑controlled authoring to bind evidence to each input revision and guide bounded repair before acceptance. Across 200 prompts, MOOSEnger raises executable success from 5% to 89.5% with GPT‑5.2 and from 0% to 76.5% with Gemma 4 31B, and a ten‑case benchmark shows all generated inputs meet semantic alignment, with eight also meeting numerical‑accuracy criteria.

By Mengnan Li, Jason Miller, Zaid Abulawi, Zachary Prince, Matt Kohl, Jack M. Cavaluzzi, Guillaume Giudicelli, Casey T. Icenhour, Alexander Lindsay, Cody Permann
arXiv AI
Aug 18

AeroCopilotBench: A Two-Tier Benchmark for Evaluating LLM Agents as Aviation Copilots in an Interactive Virtual Cockpit Environment

arXiv:2608. 16349v1 Announce Type: new Abstract: Large language model (LLM) agents may assist flight crews with complex decisions and task execution, but existing aviation evaluations centered on static knowledge do not support systematic testing of procedural execution and safety compliance in interactive environments.

By Yuchen Yuan, Zhenghuang Wu, Yuangan Li, Liang Ma, Ke Li
arXiv AI
Jul 22

SENTINEL: A Multi-Level Formal Framework for Safety Evaluation of Foundation Model-based Embodied Agents

arXiv:2510. 12985v3 Announce Type: replace Abstract: We present SENTINEL, a framework for formally evaluating the physical safety of foundation model (FM)-based embodied agents.

By Simon Sinong Zhan, Philip Wang, Yao Liu, Yiyan Peng, Zinan Wang, Qineng Wang, Zhian Ruan, Xiangyu Shi, Xinyu Cao, Frank Yang, Zhenyang Ni, Kangrui Wang, Ruohan Zhang, Huajie Shao, Manling Li, Qi Zhu
arXiv AI
Jun 17

Blueprint First, Model Second: A Framework for Deterministic LLM Workflow

arXiv:2508. 02721v2 Announce Type: replace-cross Abstract: While powerful, the inherent non-determinism of large language model (LLM) agents limits their application in structured operational environments where procedural fidelity and predictable execution are strict requirements.

By Libin Qiu, Yuhang Ye, Zhirong Gao, Xide Zou, Junfu Chen, Ziming Gui, Weizhi Huang, Xiaobo Xue, Wenkai Qiu, Kun Zhao