arXiv:2606. 03061v1 Announce Type: cross Abstract: Emerging distributed computing paradigms, such as the computing continuum, are inherently heterogeneous, stochastic, and complex.
By Alfreds Lapkovskis, Ali Beikmohammadi, Sindri Magn\'usson, Praveen Kumar Donta
arXiv:2607. 26602v1 Announce Type: cross Abstract: The rapid development of the Internet of Everything (IoE) is accelerating the adoption of intelligent applications.
By Haijun Zhang, Zhuojun Duan, Zijun Wu, Xu Ma, Yuzheng Ren
arXiv:2609.17193v1 Announce Type: new
Abstract: Large language model (LLM)-powered agentic AI services increasingly demand low-latency inference, motivating the deployment of LLMs across distributed...
By Zhen Li, Jun Cai, Haoran Gao, An Li, Tan Li
arXiv:2607. 18288v1 Announce Type: new Abstract: Load imbalance across edge and cloud layers degrades latency performance in hierarchical edge-cloud computing (HECC) systems under dynamic task arrivals and heterogeneous resources, leading to severe queuing delays and inefficient resource utilization.
By Vo Phi Son, Van-Dinh Nguyen, Ngoc Hung Nguyen, Trinh Van Chien, Symeon Chatzinotas
MCRL2 is a reinforcement learning framework that enhances microservice scheduling in cloud data centers by integrating multi-resource cross-attention-based representation learning. It introduces MCRL, a representation learning component that captures structured interactions among nodes, resources, and microservices, and couples this with an actor‑critic architecture and a maximum entropy objective. Experiments on real production cluster traces show that MCRL2 outperforms existing baselines in load balancing, scheduling success rate, and average completion time across diverse workloads.
By Tiangang Li, Shi Ying, Xiangbo Tian, Chuan Shi, Ding Xiao
The paper proposes an inference auction for large language model (LLM) APIs, enabling users to bid for faster service when compute demand exceeds capacity. The auction allocates priority efficiently without increasing latency, and includes fast algorithms for truthful bidding and an autobidding agent that adjusts bids within a user’s budget to maximize utility. Experiments show the auction improves system welfare while preserving the cache utilization and latency benefits of the SGLang inference framework.
By Keegan Harris, Siddharth Prasad, Asher Trockman, Nika Haghtalab, Michael I. Jordan
arXiv:2609.37085v1 Announce Type: cross
Abstract: Data-intensive services in the Computing Continuum must balance analytics quality, resource usage, and cost across heterogeneous nodes with limited a...
By Javier Mateos-Bravo, Sergio Laso, Juan Luis Herrera, Ilir Murturi, Pantelis Frangoudis, Schahram Dustdar
arXiv:2606. 11440v1 Announce Type: new Abstract: Existing multi-agent LLM orchestration methods, ranging from brute-force ensembles to learned routers, select models and topologies based on task and model features.
By Ahasan Kabir, Jiaqi Xue, Mengxin Zheng, Qian Lou
arXiv:2608. 02031v1 Announce Type: cross Abstract: This paper investigates collaborative mobile edge computing (MEC) servers for large language model (LLM) inference under soft deadline constraints.
By Ngoc Hung Nguyen, Bjorn Landfeldt
arXiv:2605. 02965v2 Announce Type: replace Abstract: Artificial intelligence-generated content (AIGC) has emerged as a transformative paradigm for automating the creation of diverse and customized content, giving rise to rapidly growing computational workloads in cloud data centers.
By Yang Fu, Peng Qin, Liming Chen, Zihao Zhang, Hao Yu, Yifei Wang
arXiv:2508. 06133v4 Announce Type: replace-cross Abstract: We study offline scheduling for large language model (LLM) serving under a fixed KV-cache memory budget, where requests have heterogeneous prompt (prefill) and response (decode) lengths.
By Meixuan Wang, Yinyu Ye, Zijie Zhou
DART-FL is a multitask federated learning framework designed for edge devices that must balance online inference and model training under limited resources. It dynamically allocates resources between inference and training based on current inference backlog and service capacity, then distributes remaining training capacity among tasks using a queue‑aware scheduler that adjusts loss weights. Experiments on image classification datasets with synthetic and real workloads show that DART‑FL adapts to bursty inference demand, improving accuracy for high‑demand tasks while preserving overall multitask performance.
By Yiming Xie, Pinrui Yu, Geng Yuan, Xue Lin, Ningfang Mi