OpenAI Blog

Scaling Kubernetes to 2,500 nodes

OpenAI Blog
Jan 25, 2021

Scaling Kubernetes to 7,500 nodes

We’ve scaled Kubernetes clusters to 7,500 nodes, producing a scalable infrastructure for large models like GPT-3, CLIP, and DALL·E, but also for rapid small-scale iterative research such as Scaling Laws for Neural Language Models.

arXiv Machine Learning
Aug 31

Agentic-Kube: A Graph-Enhanced Multi-Agent Reinforcement Learning Framework for Multi-Objective Kubernetes Scheduling

Agentic‑Kube is a cooperative multi‑agent reinforcement learning framework for Kubernetes pod placement that splits the multi‑objective scheduling problem into cost minimisation, anti‑affinity fault tolerance, and vector resource balancing, each handled by a dedicated sub‑agent. It uses a bipartite Graph Convolutional Network to model host‑pod dependencies, a two‑stage monotonic QMIX value factorisation network for joint action coherence, and a plurality voting consensus with action feasibility masking. Evaluations on Google Kubernetes Engine and large‑scale clusters show Pareto‑efficient placements, a 53% reduction in anti‑affinity collisions, a 65% spot instance allocation ratio, and sub‑30 ms decision latencies up to 1,000 nodes without container restarts.

By Hamed Hamzeh
arXiv Machine Learning
Sep 11

Composable CXL Memory as a Kubernetes-Native Shared Memory for LLM Serving

The paper introduces a Kubernetes Dynamic Resource Allocation driver that treats composable CXL memory as a schedulable cluster resource, enabling cross-node shared memory for large language model (LLM) serving. By composing CXL regions on demand, materializing them as DAX devices, and exposing them via a single Container Device Interface name, pods on different nodes can access the same physical memory region. A shared‑memory connector for vLLM/llm‑d uses this region as a KV‑cache tier, eliminating external metadata services and achieving significant reductions in time‑to‑first‑token (TTFT) with minimal additional latency compared to same‑node reuse.

By Hongjian Fan, Kevin Zhang, David Habinsky, Sean Dykstra
arXiv Computation and Language
Sep 21

Towards Secure Cloud-Native Computing: Unveiling Kubernetes Misconfigurations with Large Language Models

The paper investigates how Large Language Models can help detect misconfigurations in Kubernetes, the dominant platform for orchestrating containerized applications in cloud‑native environments. It presents a taxonomy of common misconfiguration types, evaluates existing detection tools, and analyzes which Kubernetes objects are most vulnerable and how severe the issues can be. The study demonstrates that advanced machine learning, particularly LLMs, can offer new insights and improve the effectiveness of misconfiguration detection methods.

By Mostafa Anouar Ghorab, Mohamed Aymen Saied
arXiv AI
Sep 24

Crossflow: Prefill-Decode Elasticity for Agentic LLM Serving

Crossflow introduces an elastic boundary for prefilling and decoding in large language model serving, allowing decode nodes to publish short‑lived leases that limit prefilling resources and output projections. By adapting to dynamic phase demand, Crossflow improves token throughput by 16.2‑17.4% on average and up to 43.4% under high load, while consistently reducing mean time‑to‑first‑token. The approach eliminates the inefficiencies of static partitioning, which can leave 17% of cluster capacity idle or cause queueing and lost throughput.

By Yi Xu, Ehsan K. Ardestani, Wenyin Fu, Martin Schatz, Krishna Malladi, Zhan Shu, Adnan Aziz, Shobhit Kanaujia, Ajit Mathews, Chunqiang Tang
arXiv Machine Learning
Jul 27

Cloud-Native Evaluation-as-a-Service: A Microservices Architecture for Scalable AI Monitoring with Conformal Guarantees

arXiv:2607. 21623v1 Announce Type: new Abstract: We present EaaS, a cloud-native reference architecture that operationalizes AI evaluation methods as six stateless Kubernetes microservices: conformal prediction with finite-sample-corrected Adaptive Prediction Sets, calibration assessment, drift detection via RFF-approximated Maximum Mean Discrepancy, fairness monitoring with bootstrap confidence intervals, a DAG-based pipeline orchestrator, and a result storage API.

By Lei Yang