Large language model (LLM)-assisted energy-management tools can translate natural-language context into structured grid commands, but syntactic validity does not imply physical admissibility. This pap...
arXiv:2609.22476v1 Announce Type: cross
Abstract: Transmission system operators face rising complexity from renewable integration, reduced inertia, and tighter security margins. Large language models...
By Costas Mylonas, Magda Foti, Emmanouel Varvarigos
arXiv:2607. 18147v1 Announce Type: cross Abstract: Large language models (LLMs) and agentic AI systems have evolved from natural language tasks to using external tools to plan, retrieve, and act in technical domains.
By Daniela Rojas, Abdulwahab Albassam, Aidan G. Leung, Jett Ngo, Ryan Luo, Peter R. Quawas, Junpyung Kim, Kangkai Liang, Mansi Nanavati, Jonathan Mai, Meng-Chi Tsai, Yun-Tong Tsai, Yize Chen, Yuanyuan Shi
The paper introduces PACE (Policy‑Attested Contract Execution), a framework that sits between large‑language‑model (LLM) based autonomous AI agents and on‑chain DeFi operations. PACE defines typed transaction intents, a deterministic policy verifier, and signed Policy Decision Records (PDRs) that cryptographically bind an approved intent, policy, and simulation report to the exact on‑chain execution bytes, providing replay and expiration protection. In evaluations across 40 tasks and six baselines, PACE achieves zero unsafe executions and zero false positives, outperforming unguarded agents by a large margin.
By Rabimba Karanjai (Larry), Yang Lu (Larry), Richard Williamson (Larry), Hemanth Hm (Larry), Prakhar Mehrotra (Larry), Lei Xu (Larry), Weidong (Larry), Shi
PLCBench is a hardware‑in‑the‑loop framework that evaluates whether autonomous large language model agents can transform network‑reachable programmable logic controllers (PLCs) into sustained physical threats. It integrates vendor‑native PLC interaction, commercial PLC execution, closed‑loop process simulation, and deterministic diagnostics to classify episodes into usable interaction, process‑linked manipulation, and sustained physical impact. Across 240 real‑PLC episodes with five LLM families, 31.3% achieved sustained physical objectives, revealing that richer process observation improves success rates and pinpointing failure points for future defense research.
By Yitian Zhou, Jingyu Zheng, Qiliang Jiang, Linkang Du, Haoming Liu, Lichao Wu, Shiyi Zhao, Mengxiang Liu, Ruilong Deng
arXiv:2606. 20950v2 Announce Type: replace Abstract: Executable evaluation -- checking the consequences of an agent's actions with a program rather than grading its prose -- has become a prominent way to assess tool-using AI agents in software settings.
By Sergei Trashchenkov
arXiv:2605. 17909v2 Announce Type: replace Abstract: As autonomous agentic systems scale across regulated critical infrastructures, the lack of mechanistic, hardware-rooted enforcement for high-frequency policy updates presents a fundamental safety gap.
By Riddhi Mohan Sharma
arXiv:2606. 09549v1 Announce Type: cross Abstract: Tool-using large language model (LLM) agents face two distinct security failures: unauthorized external actions and exposure of sensitive plaintext inside the runtime before any final output check can intervene.
By Yuhan Ma, Stefan Schmid
Grid‑Orch is a framework that connects Large Language Models (LLMs) to power system simulation via the Model Context Protocol (MCP), allowing engineers to conduct complex distribution grid analyses using natural language. It offers 36 domain‑specific tools across eleven categories—including power flow, voltage analysis, quasi‑static time‑series simulation, and automated optimization—implemented with OpenDSS as the reference engine. The platform supports both cloud‑hosted and locally deployed LLMs, enabling air‑gapped operation, and demonstrates that tasks such as DER interconnection screening can be completed in under two minutes with results identical to traditional scripting.
By Boming Liu, Jin Dong, Jianming Lian
The paper proposes a split‑control architecture for adaptive security at the network edge, where an untrusted planner emits typed security intents that are vetted by a deterministic governor before being enacted. The governor checks each intent against safety, resource, temporal‑stability, and proportionality invariants, issuing signed receipts for admitted actions that are compiled into eBPF map updates. Experiments on a Raspberry Pi 5 connected to a university 5G test network show the governor can admit, reject, and bound intents at microsecond cost without disrupting protected‑flow regularity.
By Ijaz Ahmad, Ijaz Ahmad, Flavio Esposito, Erkki Harjula
arXiv:2609.26048v1 Announce Type: cross
Abstract: Language-model agents often reach a working solution and then fail to consistently deliver it. We study runtime policies: targeted natural-language i...
By Nikita Agarwal, Nivedit Jain
arXiv:2609.16313v1 Announce Type: cross
Abstract: In agentic distributed systems, an agent may be authorized to mutate external infrastructure while lacking evidence that the mutation is ready to exe...
By Jun He, Deying Yu