arXiv Machine Learning

When Does Learning Beat Heuristics? A Case Study in Kubernetes Scheduler Score Plugins

The study investigates whether learned scoring functions can outperform hand‑tuned heuristics in Kubernetes node‑selection. Two models—a Random Forest on engineered features and a graph neural network on job dependency graphs—were trained on a large production trace; both achieved modest regression gains (R²≈0.042) but lagged behind a simple free‑CPU heuristic in Top‑1 ranking accuracy (65‑66% vs. 74‑84%). The authors attribute this gap to objective mismatch, noting that pointwise regression rather than a ranking‑specific loss likely limits performance.

arXiv Machine Learning
Sep 22

Efficient K-generalizable Learned Search

arXiv:2603.06159v2 Announce Type: replace-cross Abstract: Learned top-K search improves the accuracy-latency trade-off of graph-based vector search, but existing methods are designed for a fixed K: s...

By Yifan Peng, Jiafei Fan, Xingda Wei, Sijie Shen, Rong Chen, Jianning Wang, Xiaojian Luo, Wenyuan Yu, Jingren Zhou, Haibo Chen
arXiv Machine Learning
Aug 31

Agentic-Kube: A Graph-Enhanced Multi-Agent Reinforcement Learning Framework for Multi-Objective Kubernetes Scheduling

Agentic‑Kube is a cooperative multi‑agent reinforcement learning framework for Kubernetes pod placement that splits the multi‑objective scheduling problem into cost minimisation, anti‑affinity fault tolerance, and vector resource balancing, each handled by a dedicated sub‑agent. It uses a bipartite Graph Convolutional Network to model host‑pod dependencies, a two‑stage monotonic QMIX value factorisation network for joint action coherence, and a plurality voting consensus with action feasibility masking. Evaluations on Google Kubernetes Engine and large‑scale clusters show Pareto‑efficient placements, a 53% reduction in anti‑affinity collisions, a 65% spot instance allocation ratio, and sub‑30 ms decision latencies up to 1,000 nodes without container restarts.

By Hamed Hamzeh
arXiv AI
Aug 24

UpgradeBench: A Decision-Centric Benchmark for Upgrading Fine-Tuned LLM Specialists

UpgradeBench is a decision‑centric longitudinal benchmark that evaluates how fine‑tuned language‑model specialists should be handled when new base‑model releases occur. It covers four consecutive Qwen releases, a continuation checkpoint, six tasks, two model sizes, and OLMo checkpoints with known training lineage, and examines whether retraining, adapter transfer, or other recovery strategies improve specialist performance. The benchmark reveals that upgrade gains vary by task and release interval, that direct adapter copying is sensitive to pretraining distance, and that teacher relabeling can recover specialists without new annotations. "whyItMatters":"The study provides actionable insights into the cost‑effective management of specialist models across model releases, showing how to balance retraining effort with performance gains."

By Ye Chen, Weining Zhang