arXiv AI By Youssef Ennouri, Soonhoi Ha

MeanField Surrogate Modeling for Scalable Runtime Scheduling of Concurrent Heterogeneous AI Inference on Shared GPUs

Read the original on arXiv AI →

The paper introduces a MeanField surrogate model for predicting performance of concurrent heterogeneous AI inference workloads on shared GPUs, reducing profiling complexity from combinatorial to linear in the number of models. Experiments with up to six models show high accuracy (R²≈0.96) and efficient integration into a genetic algorithm scheduler, achieving near-exhaustive search performance with minimal runtime overhead.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 19

KernelArc: A Multi-Agent Framework for GPU Kernel Optimization

KernelArc is a multi-agent framework designed to autonomously optimize GPU kernels across diverse workloads. It employs strategy-specialized agents that run concurrently, coordinating via conclusions-only shared memory, a deterministic benchmark guard, and read-only cross-agent state with plateau-triggered drafting. Evaluated on NVIDIA H100 and B200 GPUs with SOL-ExecBench workloads, KernelArc produced top-ranked implementations for tasks such as BF16 GEMM, cuBLASLt configuration tables, and various attention mechanisms, achieving first place on several leaderboard categories.

By Joyjit Kundu, Ben Stoffelen, Kaili Wang, Peter Vrancx, Ludovic Denoyer
Hugging Face Trending Papers
5d ago

AgentPerfBench: A Benchmarking and Evaluation Suite for Inference Performance of Agentic LLMs

AgentPerfBench is a new benchmarking suite designed to evaluate the inference performance of agentic large language models (LLMs) that handle multi‑turn, tool‑using, and context‑expanding tasks. It builds on real traces from agentic benchmarks such as SWE‑Bench and TerminalBench, and generates synthetic profiles that reflect realistic input/output lengths and turn counts. The suite also provides kernel‑level Nsight Compute traces and a multi‑dimensional roofline model to identify hardware bottlenecks and quantify the gap between traditional chat benchmarks and agentic workloads.