Large language model (LLM)-assisted energy-management tools can translate natural-language context into structured grid commands, but syntactic validity does not imply physical admissibility. This pap...
arXiv:2609.22476v1 Announce Type: cross
Abstract: Transmission system operators face rising complexity from renewable integration, reduced inertia, and tighter security margins. Large language models...
By Costas Mylonas, Magda Foti, Emmanouel Varvarigos
arXiv:2607. 18147v1 Announce Type: cross Abstract: Large language models (LLMs) and agentic AI systems have evolved from natural language tasks to using external tools to plan, retrieve, and act in technical domains.
By Daniela Rojas, Abdulwahab Albassam, Aidan G. Leung, Jett Ngo, Ryan Luo, Peter R. Quawas, Junpyung Kim, Kangkai Liang, Mansi Nanavati, Jonathan Mai, Meng-Chi Tsai, Yun-Tong Tsai, Yize Chen, Yuanyuan Shi
The paper introduces PACE (Policy‑Attested Contract Execution), a framework that sits between large‑language‑model (LLM) based autonomous AI agents and on‑chain DeFi operations. PACE defines typed transaction intents, a deterministic policy verifier, and signed Policy Decision Records (PDRs) that cryptographically bind an approved intent, policy, and simulation report to the exact on‑chain execution bytes, providing replay and expiration protection. In evaluations across 40 tasks and six baselines, PACE achieves zero unsafe executions and zero false positives, outperforming unguarded agents by a large margin.
By Rabimba Karanjai (Larry), Yang Lu (Larry), Richard Williamson (Larry), Hemanth Hm (Larry), Prakhar Mehrotra (Larry), Lei Xu (Larry), Weidong (Larry), Shi
PLCBench is a hardware‑in‑the‑loop framework that evaluates whether autonomous large language model agents can transform network‑reachable programmable logic controllers (PLCs) into sustained physical threats. It integrates vendor‑native PLC interaction, commercial PLC execution, closed‑loop process simulation, and deterministic diagnostics to classify episodes into usable interaction, process‑linked manipulation, and sustained physical impact. Across 240 real‑PLC episodes with five LLM families, 31.3% achieved sustained physical objectives, revealing that richer process observation improves success rates and pinpointing failure points for future defense research.
By Yitian Zhou, Jingyu Zheng, Qiliang Jiang, Linkang Du, Haoming Liu, Lichao Wu, Shiyi Zhao, Mengxiang Liu, Ruilong Deng
arXiv:2606. 20950v2 Announce Type: replace Abstract: Executable evaluation -- checking the consequences of an agent's actions with a program rather than grading its prose -- has become a prominent way to assess tool-using AI agents in software settings.
By Sergei Trashchenkov