arXiv Machine Learning By Wang Xuying, Zhibek Sarypbekova

When Does Learning Beat Heuristics? A Case Study in Kubernetes Scheduler Score Plugins

Read the original on arXiv Machine Learning →

The study investigates whether learned scoring functions can outperform hand‑tuned heuristics in Kubernetes node‑selection. Two models—a Random Forest on engineered features and a graph neural network on job dependency graphs—were trained on a large production trace; both achieved modest regression gains (R²≈0.042) but lagged behind a simple free‑CPU heuristic in Top‑1 ranking accuracy (65‑66% vs. 74‑84%). The authors attribute this gap to objective mismatch, noting that pointwise regression rather than a ranking‑specific loss likely limits performance.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 22

Efficient K-generalizable Learned Search

arXiv:2603.06159v2 Announce Type: replace-cross Abstract: Learned top-K search improves the accuracy-latency trade-off of graph-based vector search, but existing methods are designed for a fixed K: s...

By Yifan Peng, Jiafei Fan, Xingda Wei, Sijie Shen, Rong Chen, Jianning Wang, Xiaojian Luo, Wenyuan Yu, Jingren Zhou, Haibo Chen
arXiv Machine Learning
Aug 31

Agentic-Kube: A Graph-Enhanced Multi-Agent Reinforcement Learning Framework for Multi-Objective Kubernetes Scheduling

Agentic‑Kube is a cooperative multi‑agent reinforcement learning framework for Kubernetes pod placement that splits the multi‑objective scheduling problem into cost minimisation, anti‑affinity fault tolerance, and vector resource balancing, each handled by a dedicated sub‑agent. It uses a bipartite Graph Convolutional Network to model host‑pod dependencies, a two‑stage monotonic QMIX value factorisation network for joint action coherence, and a plurality voting consensus with action feasibility masking. Evaluations on Google Kubernetes Engine and large‑scale clusters show Pareto‑efficient placements, a 53% reduction in anti‑affinity collisions, a 65% spot instance allocation ratio, and sub‑30 ms decision latencies up to 1,000 nodes without container restarts.

By Hamed Hamzeh