arXiv:2607. 19349v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as always-on online services, making efficient LLM serving a critical systems challenge.
By Tiancheng Zhang, Shaoyuan Huang, Mingyuan Wang, Yunfeng Zhao, Xiaofei Wang, Wenyu Wang
The paper introduces agentic-eCAL, an extension of the Energy Cost of AI Lifecycle metric to evaluate multi‑agent AI workflows across the edge‑cloud continuum. By combining a two‑rate energy model with OSI‑layer transport analysis, the authors quantify that inter‑agent text transfer accounts for only 0.25% of total workflow energy, highlighting that the main energy cost lies in additional inference and context processing triggered by communication. The study uses extensive GPU benchmarks on NVIDIA A100/H100 with 16 open‑weight models and 8 orchestration topologies to validate the metric and explore placement implications.
By Carolina Fortuna, Vid Han\v{z}el, Tim Strnad, Bla\v{z} Bertalani\v{c}
arXiv:2609.35916v1 Announce Type: cross
Abstract: Real-world embodied agents often pursue independent objectives within a shared physical environment, where their actions can alter the conditions fac...
By Jie Yang, Jiajun Chen, Jiazheng Zhou, Mianqiu Huang, Yining Zheng, Yuxin Wang, Xipeng Qiu
arXiv:2608. 00101v1 Announce Type: cross Abstract: AI coding agents like GitHub Copilot, Claude Code, and Codex interleave multi-step LLM inference with tool execution, creating a workload different from chatbots.
By Banruo Liu, Haoran Qiu, \'I\~nigo Goiri, Rodrigo Fonseca, Ricardo Bianchini, Esha Choukse
arXiv:2606. 31209v1 Announce Type: new Abstract: Interactive traffic simulation is a vital world model for autonomous driving.
By Lingyu Xiao, Zexin Feng, Xintao Yan
arXiv:2608. 15127v1 Announce Type: cross Abstract: Agentic applications are shifting AI serving from isolated model inference to long-running workloads in which LLMs coordinate tools, environments, and persistent state.
By Chaokun Chang, Yukun Zhou, Kaihua Fu, Dakai An, Tianyu Feng, Hanfeng Lu, Sheng Yao, Pu Guo, Yinghao Yu, Yizhou Shan, Bo Li, Binhang Yuan, Wei Wang
arXiv:2608. 01725v1 Announce Type: cross Abstract: Modern computing and networking infrastructure emits telemetry continuously, yet operators convert it into decisions with a separate predictor per task, entity, and horizon.
By Zifan Zhang, Zhichao Hou, Tingxiang Ji, Yuchen Liu
arXiv:2606. 11440v1 Announce Type: new Abstract: Existing multi-agent LLM orchestration methods, ranging from brute-force ensembles to learned routers, select models and topologies based on task and model features.
By Ahasan Kabir, Jiaqi Xue, Mengxin Zheng, Qian Lou
arXiv:2607. 16200v1 Announce Type: new Abstract: AI agent systems that couple large language models (LLMs) with external tools and APIs are inherently non-deterministic: LLM sampling variance, external API state, CDN infrastructure headers, and execution-environment noise collectively prevent any prior agent run from being faithfully re-executed.
By Rasheed Mudasiru
arXiv:2604. 17456v2 Announce Type: replace Abstract: Large language model (LLM) agents have shown strong capabilities in long-horizon reasoning, tool use, and decision-making in digital environments, yet extending them to physically grounded systems remains challenging.
By Siqi Lai, Pan Zhang, Yuping Zhou, Jindong Han, Yansong Ning, Hao Liu
arXiv:2606. 10662v1 Announce Type: cross Abstract: Multi-agent systems (MAS) can scale large language model reasoning at test time by decomposing complex problems into parallel subtasks.
By Yuzhen Mao, Azalia Mirhoseini
arXiv:2608.23179v1 Announce Type: cross
Abstract: Large language model (LLM) agents are increasingly attractive for automating network configuration, yet their reliability and failure patterns are po...
By Chang Liu, Xiaohui Xie, Xinyi Chen, Yong Cui