arXiv:2606. 03077v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become a standard post-training paradigm for large language models (LLMs), extending beyond preference alignment to complex reasoning and multi-turn agentic behaviors.
By Kaiwen Chen, Xin Tan, Jingzong Li, Hong Xu
arXiv:2605. 26418v2 Announce Type: replace-cross Abstract: A properly calibrated rule-based autoscaler can beat every one of six mainstream deep reinforcement learning (DRL) algorithms on cost across every workload we test - so when, if ever, does DRL actually help?
By Guilin Zhang, Chuanyi Sun, Kai Zhao, Shahryar Sarkani, John Fossaceca
Agentic‑Kube is a cooperative multi‑agent reinforcement learning framework for Kubernetes pod placement that splits the multi‑objective scheduling problem into cost minimisation, anti‑affinity fault tolerance, and vector resource balancing, each handled by a dedicated sub‑agent. It uses a bipartite Graph Convolutional Network to model host‑pod dependencies, a two‑stage monotonic QMIX value factorisation network for joint action coherence, and a plurality voting consensus with action feasibility masking. Evaluations on Google Kubernetes Engine and large‑scale clusters show Pareto‑efficient placements, a 53% reduction in anti‑affinity collisions, a 65% spot instance allocation ratio, and sub‑30 ms decision latencies up to 1,000 nodes without container restarts.
By Hamed Hamzeh
The paper introduces CloudEdgeVLA, a cloud‑edge policy for Vision‑Language‑Action models that treats temporal misalignment as a representation‑learning problem. It encodes delayed observations into slowly varying task features on the cloud while a lightweight edge head fuses the latest cloud feature with current local vision. Experiments on four LIBERO suites show that CloudEdgeVLA retains 63.8–78.0% success under a 40‑step delay window, far outperforming VLASH and single‑path baselines.
By Daojie Peng, Fulong Ma, Bingtao Wang, Sheng Wang, Jun Ma
arXiv:2609.37085v1 Announce Type: cross
Abstract: Data-intensive services in the Computing Continuum must balance analytics quality, resource usage, and cost across heterogeneous nodes with limited a...
By Javier Mateos-Bravo, Sergio Laso, Juan Luis Herrera, Ilir Murturi, Pantelis Frangoudis, Schahram Dustdar
arXiv:2609.17193v1 Announce Type: new
Abstract: Large language model (LLM)-powered agentic AI services increasingly demand low-latency inference, motivating the deployment of LLMs across distributed...
By Zhen Li, Jun Cai, Haoran Gao, An Li, Tan Li
arXiv:2608.29255v1 Announce Type: cross
Abstract: Artificial Intelligence-Generated Content (AIGC) services employ Generative AI (GenAI) models to automatically generate diverse content. Mobile AIGC...
By Chongzhi Wu, Zhengtao Li, Jiawen Kang, Jinbo Wen, Xiaohuan Li, Maomao Zhang, Ekram Hossain
The paper introduces COMLLM, a generative framework that combines Group Relative Policy Optimization with a Look‑Ahead Collaborative Simulation to enable multi‑turn reasoning for task offloading in Mobile Edge Computing. By performing multi‑step Monte Carlo rollouts that jointly model server queue dynamics, COMLLM incorporates long‑term system evolution into its reward design, achieving near‑optimal latency and improved load‑balancing fairness. The framework demonstrates zero‑shot scalability to larger network topologies, outperforming supervised fine‑tuning, deep reinforcement learning, and heuristic baselines without requiring retraining.
By Ning Yang, Chuangxin Cheng, Haijun Zhang
arXiv:2608. 10357v1 Announce Type: cross Abstract: Long-horizon tool-using agents must reason over user goals, domain policies, tool calls, simulator state, and delayed verifiable rewards.
By Zelei Cheng, Amritansh Mishra, Sambit Sahu, William Campbell
arXiv:2606. 00756v1 Announce Type: new Abstract: Deploying lightweight Large Language Model (LLM) agents on edge servers can reduce latency and move agentic services closer to users, but resource-constrained edge models often struggle with long-horizon tasks that require persistent memory, subgoal tracking, and reflection.
By Yannan Wang, Longli Yang, Zhen Liu, Abhishek Kumar, Carsten Maple
arXiv:2512. 22560v2 Announce Type: replace-cross Abstract: Agentic Reinforcement Learning (RL) trains LLMs through multi-turn interactions with environments, producing workloads that mix compute-bound prefill, bandwidth-bound decoding, CPU-heavy environment execution, and bursty reward evaluation.
By Wei Gao, Yuheng Zhao, Tianyuan Wu, Shaopan Xiong, Weixun Wang, Dakai An, Lunxi Cao, Dilxat Muhtar, Zichen Liu, Haizhou Zhao, Ju Huang, Siran Yang, Yongbin Li, Wenbo Su, Jiamang Wang, Lin Qu, Bo Zheng, Wei Wang
arXiv:2606. 11440v1 Announce Type: new Abstract: Existing multi-agent LLM orchestration methods, ranging from brute-force ensembles to learned routers, select models and topologies based on task and model features.
By Ahasan Kabir, Jiaqi Xue, Mengxin Zheng, Qian Lou