arXiv:2607. 14158v1 Announce Type: new Abstract: This position paper explores how Agentic AI and Model Context Protocol (MCP) can support power-grid studies in a Transmission System Operator (TSO) context.
By J\'er\^ome Picault, Cl\'ement Goubet
Grid‑Orch is a framework that connects Large Language Models (LLMs) to power system simulation via the Model Context Protocol (MCP), allowing engineers to conduct complex distribution grid analyses using natural language. It offers 36 domain‑specific tools across eleven categories—including power flow, voltage analysis, quasi‑static time‑series simulation, and automated optimization—implemented with OpenDSS as the reference engine. The platform supports both cloud‑hosted and locally deployed LLMs, enabling air‑gapped operation, and demonstrates that tasks such as DER interconnection screening can be completed in under two minutes with results identical to traditional scripting.
By Boming Liu, Jin Dong, Jianming Lian
arXiv:2607. 18147v1 Announce Type: cross Abstract: Large language models (LLMs) and agentic AI systems have evolved from natural language tasks to using external tools to plan, retrieve, and act in technical domains.
By Daniela Rojas, Abdulwahab Albassam, Aidan G. Leung, Jett Ngo, Ryan Luo, Peter R. Quawas, Junpyung Kim, Kangkai Liang, Mansi Nanavati, Jonathan Mai, Meng-Chi Tsai, Yun-Tong Tsai, Yize Chen, Yuanyuan Shi
arXiv:2608.24650v1 Announce Type: cross
Abstract: System-level simulation is an essential tool for exploring the rapidly expanding design space of LLM serving systems, where real deployments remain c...
By Wonung Kim, Hyunmin Choi, Minsu Kim, Jaehong Cho, Yeongwook Kim, Jongse Park
AURORA is a natural‑language‑driven framework that treats air‑ground scenario generation as a compilation process with verification. It introduces the Air‑Ground Scenario Graph (AGSG), a typed intermediate representation linking agents, missions, events, communication, and success conditions, enabling joint grounding, temporal planning, pre‑execution checks, runtime verification, failure localization, and bounded repair. The authors also present AURORA‑Bench to evaluate not only execution but faithful realization of requested interactions, showing that structured execution and runtime verification improve reliability and that explicit intermediate representations facilitate verifiable and repairable co‑simulation.
By Keshu Wu, Hao Zhang, Rui Gan, Xiangbo Gao, Xiaopeng Li, Zhengzhong Tu, Yang Zhou
arXiv:2609.00384v1 Announce Type: new
Abstract: Large Language Model (LLM) agents increasingly automate multi-step engineering workflows through tool use, interpretation of intermediate results, and...
By Riccardo Mansutti, Andrea Pomarico, Robert Jakob, Qian Zhang, Alberto Berizzi, Kevin O'Sullivan
arXiv:2608. 08691v1 Announce Type: new Abstract: Residential virtual power plants (VPPs) can provide grid flexibility by shifting household demand, but physical flexibility becomes dependable capacity only when residents authorize a plan and the promised response is delivered.
By Xudong Wu, Zeqing Wu, Jiarui Zhang, Xuhao Fan, Ziang Ding, Yuming Zhuang, Mingqi Yuan, Yilun Du, Hongjie Jia, Yunfei Mu, Jiayu Chen
The paper reviews 66 studies on large language models (LLMs) applied to HVAC operations in building energy systems, categorizing them by application and method families and evaluating their evidence realism and deployment readiness. It finds that most work focuses on building energy modelling, with only four studies reaching pilot-level evidence and none reporting sustained operational deployment. LLMs are currently best suited as semantic and workflow layers—such as point‑name normalisation and document‑grounded operator support—rather than autonomous HVAC controllers, and future research should target field‑validated benchmarks and safe, low‑latency LLM‑MPC/RL integrations.
By Alexander Neubauer, Tianzhen Hong, Han Li, Mengbo Yu, Amin Darbandi, Yannick F\"urst, Martin Kriegel
arXiv:2608. 15041v1 Announce Type: new Abstract: Coordinating multiple interacting units in complex engineering systems is challenging when system interactions are difficult to model, operational information is heterogeneous, and low-level actions must satisfy strict constraints.
By Changhong He, Jinda Gao, Xinkuan Liu, Le Zhang, Xizi Luo, Yu Mei
arXiv:2607. 26710v1 Announce Type: new Abstract: The rapid growth of AI workloads is turning data centers into large-scale, volatile, yet spatiotemporally flexible grid loads, creating an urgent need for coordinated electricity-computing scheduling.
By Kaiwen Jiang, Siya Xu, Ziyue Zhu, Chao Yang, Anh Tuan Luu, Haoran Luo
arXiv:2604. 22748v2 Announce Type: replace Abstract: As AI systems move from generating text to accomplishing goals through sustained interaction, the ability to model environment dynamics becomes a central bottleneck.
By Meng Chu, Xuan Billy Zhang, Kevin Qinghong Lin, Lingdong Kong, Jize Zhang, Teng Tu, Weijian Ma, Ziqi Huang, Senqiao Yang, Wei Huang, Yeying Jin, Zhefan Rao, Jinhui Ye, Xinyu Lin, Xichen Zhang, Qisheng Hu, Shuai Yang, Leyang Shen, Wei Chow, Yifei Dong, Fengyi Wu, Quanyu Long, Bin Xia, Shaozuo Yu, Mingkang Zhu, Wenhu Zhang, Jiehui Huang, Haokun Gui, Runyi Li, Shiyi Du, Xu Huang, Dong Huang, Rui Liu, Chenyu Tang, Xuhang Chen, Chengzu Li, Haoxuan Che, Long Chen, Qifeng Chen, Wenxuan Zhang, Wenya Wang, Xiaojuan Qi, Yang Deng, Yanwei Li, Mike Zheng Shou, Zhi-Qi Cheng, See-Kiong Ng, Ziwei Liu, Philip Torr, Jiaya Jia
The paper introduces Pufibara, an agent harness designed to maintain engineering state and evidence across revisions in Modelica-based physical system modeling. It also presents a 232-task Modelica Agent Workflow Benchmark covering model repair, generation, and tuning, evaluated by an external benchmark-owned evaluator. Experiments show Pufibara outperforms Claude Code in task success and resource efficiency across two LLM backends.
By Zizhe Wang