arXiv:2608. 03609v1 Announce Type: new Abstract: Agentic systems driven by large language models (LLMs) are increasingly deployed in real-world workflows where they act on persistent operational data.
By Alejandro J. Mercado, Alessio Lomuscio
arXiv:2606. 06523v1 Announce Type: new Abstract: Equipping Large Language Models (LLMs) to execute reliable multi-step workflows has become a central challenge in artificial intelligence.
By Ruida Wang, Jerry Huang, Pengcheng Wang, Xuanqing Liu, Luyang Kong, Tong Zhang
arXiv:2606. 14790v1 Announce Type: cross Abstract: LLM-based multi-agent systems increasingly coordinate planning, reasoning, tool use, and human interaction, yet their reliability remains limited.
By Hanqi Li, Jing Peng, Zijian Wang, Lu Chen, Kai Yu
AURORA is a natural‑language‑driven framework that treats air‑ground scenario generation as a compilation process with verification. It introduces the Air‑Ground Scenario Graph (AGSG), a typed intermediate representation linking agents, missions, events, communication, and success conditions, enabling joint grounding, temporal planning, pre‑execution checks, runtime verification, failure localization, and bounded repair. The authors also present AURORA‑Bench to evaluate not only execution but faithful realization of requested interactions, showing that structured execution and runtime verification improve reliability and that explicit intermediate representations facilitate verifiable and repairable co‑simulation.
By Keshu Wu, Hao Zhang, Rui Gan, Xiangbo Gao, Xiaopeng Li, Zhengzhong Tu, Yang Zhou
arXiv:2607. 17780v1 Announce Type: cross Abstract: ETAS is a programming language for agent systems that treats model-backed agents, tool calls, prompts, typed memory, human approvals, policies, and execution traces as semantic program elements rather than library conventions.
By Huiri Tan, Yikun Wang, Puyang Zhang, Shangyu Li, Jiasi Shen
arXiv:2604. 11556v2 Announce Type: replace-cross Abstract: LLM-assisted software development has become increasingly prevalent, and can generate large-scale systems, such as compilers.
By Haoran Ding, Zhaoguo Wang, Haibo Chen
arXiv:2606. 16010v1 Announce Type: cross Abstract: Large language models have achieved impressive performance on reasoning tasks spanning mathematics, science, programming, and commonsense inference.
By Raghu Anantharangachar
The paper introduces control‑data flow separation to improve prompt optimization in multi‑agent large language model systems. By representing execution protocols as typed, validated program objects and keeping task‑relevant content as unstructured language, the method prevents prompt edits from corrupting critical routing, formatting, or termination signals. Experiments on synthetic reasoning, collaborative review generation, and insurance rating workflows show that this approach maintains 100% protocol validity while consistently enhancing task performance.
By Wentao Zhang, Syed Shariyar Murtaza, Junaid Ahmad Bhatti, Utkarsh Soni, Yifan Nie, Eugene Wen, Yuntian Deng
arXiv:2605. 09045v2 Announce Type: replace Abstract: Agentic frameworks are the software layer through which AI agents act in the world.
By Royce Moon, Lav R. Varshney
The paper introduces MAGS, a multi-agent framework that automatically generates executable programs with formal safety guarantees. MAGS translates LLM-generated code into the verification-aware language Dafny, repairs any safety violations using verifier feedback, and then compiles the verified code back into executable form. Evaluations on 220 diverse examples—including CUDA kernels, terminal scripts, and robotic-arm tasks—show a 100% success rate in producing programs that meet frozen safety specifications, with additional safety and functional tests confirming strong performance across domains.
By Albert Wu, Nicholas Roberts, Tzu-Heng Huang, Haoran Lin, Gil Friedman, Sungjun Cho, Gabriel Orlanski, Frederic Sala
arXiv:2606. 00220v1 Announce Type: cross Abstract: Formal methods provide rigorous accounts of program behavior, but practical software engineering often works through executable libraries, tests, and incremental design.
By Eric Liang
arXiv:2510. 12985v3 Announce Type: replace Abstract: We present SENTINEL, a framework for formally evaluating the physical safety of foundation model (FM)-based embodied agents.
By Simon Sinong Zhan, Philip Wang, Yao Liu, Yiyan Peng, Zinan Wang, Qineng Wang, Zhian Ruan, Xiangyu Shi, Xinyu Cao, Frank Yang, Zhenyang Ni, Kangrui Wang, Ruohan Zhang, Huajie Shao, Manling Li, Qi Zhu