Hugging Face Trending Papers

Adaptive Inference Batching using Policy Gradients

Read the original on Hugging Face Trending Papers →

Inference serving systems must balance throughput and latency under bursty, heterogeneous workloads, yet the industry standard remains static batching policies that require manual tuning and cannot adapt to shifting traffic. We investigate whether reinforcement learning (RL) can learn adaptive batching and routing policies that outperform these heuristics, training REINFORCE and PPO agents on a discrete-event simulator validated against queuing theory and production traces (Azure Functions, BurstGPT).

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv AI
Sep 4

From Prior-Guided Heuristics to Deployable Agents: Accelerating Demonstration-Driven Reinforcement Learning for Deadline-Constrained Network Control

The paper proposes a deployment‑focused framework for deadline‑constrained network control, introducing the Effective Congestion (EC) metric family and Uniform Path Grouping (UPG) heuristic to better capture traffic urgency and balance load. It integrates these with a Multi‑Agent Deep Reinforcement Learning architecture (MADRL EC (p*)) that combines a distributed scheduler and a centralized RL router. A unified training objective merges live‑reward, pre‑collected‑reward, and policy‑imitation terms, leading to the Model‑Guided Annealed Reinforcement Learning (MGA‑RL) protocol built on DDPG, which generalizes offline‑to‑online learning for demonstration‑driven training.

By Vincenzo Norman Vitale, Mohammad Solki, Antonia Maria Tulino, Andreas F. Molisch, Jaime Llorca
arXiv AI
Jun 16

RollArt: Disaggregated Multi-Task Agentic RL Training at Scale

arXiv:2512. 22560v2 Announce Type: replace-cross Abstract: Agentic Reinforcement Learning (RL) trains LLMs through multi-turn interactions with environments, producing workloads that mix compute-bound prefill, bandwidth-bound decoding, CPU-heavy environment execution, and bursty reward evaluation.

By Wei Gao, Yuheng Zhao, Tianyuan Wu, Shaopan Xiong, Weixun Wang, Dakai An, Lunxi Cao, Dilxat Muhtar, Zichen Liu, Haizhou Zhao, Ju Huang, Siran Yang, Yongbin Li, Wenbo Su, Jiamang Wang, Lin Qu, Bo Zheng, Wei Wang