arXiv Machine Learning

Design and Scheduling of an AI-based Queueing System

arXiv:2406. 06855v3 Announce Type: replace-cross Abstract: To leverage prediction models to make optimal scheduling decisions in service systems, we must understand how predictive errors impact congestion due to externalities on the delay of other jobs.

arXiv AI
Aug 28

A Multi-Modal AI Framework for Real-Time Queue Prediction, Management and Optimisation in Intelligent Border Control Systems

The paper proposes a multi‑modal AI framework for real‑time queue prediction, management, and resource optimisation in border control systems. It integrates heterogeneous data sources using LSTM networks for forecasting and applies Model Predictive Control and scheduling optimisation to generate actionable policies for officers. Evaluation on synthetic traffic data shows up to 35% reduction in prediction error, 30% lower average waiting time, and nearly 20% higher throughput compared to ARIMA and rule‑based methods.

By Varvara Mama, Eleni Veroni, Nikolaos Kapsalis, Christos D. Nikolopoulos, Anargyros T. Baklezos
arXiv Machine Learning
Aug 27

Scorpio: Serving Right Requests at the Right Time for Heterogeneous SLOs in LLM Inference

Scorpio is an LLM serving system that optimizes for heterogeneous Service Level Objectives (SLOs) such as Time to First Token (TTFT) and Time Per Output Token (TPOT). It uses adaptive scheduling across admission control, queue management, and batch selection, featuring a TTFT Guard that reorders requests by least-deadline-first and rejects unattainable ones, and a TPOT Guard that employs VBS-based admission control and a credit-based batching mechanism. Predictive modules support both guards, and evaluations show Scorpio can increase system goodput by up to 14.4× and improve SLO adherence by up to 46.5% under high load compared to state-of-the-art baselines.

By Yinghao Tang, Tingfeng Lan, Bo Pan, Xiuqi Huang, Hui Lu, Wei Chen
arXiv AI
Jun 29

Ranking Before Serving: Low-Latency LLM Serving via Pairwise Learning-to-Rank

arXiv:2510. 03243v3 Announce Type: replace-cross Abstract: Efficient scheduling of large language model (LLM) inference tasks is critical for achieving low latency and high throughput, a challenge that is becoming increasingly acute with the rise of reasoning-capable LLMs whose generation lengths are highly variable.

By Yiheng Tao, Yihe Zhang, Matthew Dearing, Xin Wang, Yuping Fan, Michael E. Papka, Zhiling Lan
arXiv AI
Jun 11

Offline Diffusion Policy for Multi-User Delay-Constrained Scheduling

arXiv:2501. 12942v2 Announce Type: replace Abstract: Effective multi-user delay-constrained scheduling is crucial in various real-world applications, including embodied AI, instant messaging, live streaming, and data center management, where efficient resource allocation is required among users with diverse delay sensitivities.

By Zhuoran Li, Ruishuo Chen, Hai Zhong, Longbo Huang