arXiv AI By Keyvan Dadashzadeh, Yuehong Zhou, Minyu Cui, Miquel Pericas

T-CCL: Resource Efficient and Performant Collective Communication using Tensor Memory Accelerator

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv AI
Sep 15

mKernel: Fast Multi-GPU, Multi-Node Fused Kernels

arXiv:2609.13585v1 Announce Type: cross Abstract: Communication has become a bottleneck in distributed training and inference of large models. Overlapping communication with computation at the granul...

By Ziming Mao, Yihan Zhang, Shawn Wei Chew, Shuang Ma, Costin Raiciu, Yang Zhou, Scott Shenker, Ion Stoica
arXiv Machine Learning
6d ago

Efficient Expert-Parallel Communication on PCIe-Connected Consumer GPUs

ThunderEP is a new communication design for expert parallelism on PCIe-connected consumer GPUs that eliminates relay hops, uses DMA engines to avoid GPU compute contention, and reduces CPU polling overhead. Integrated into vLLM, it outperforms NCCL on RTX 4090 and RTX 5090 GPUs, delivering average speedups of 2.00× for dispatch, 1.53× for combine, and up to 1.66× end‑to‑end over existing MoE inference frameworks.

By Jaehwan Lee, Sangmin Lee, Chaewon Kim, Junsik Shin, Jaejin Lee