← Back to all news
Hugging Face Blog August 8, 2025

Accelerate ND-Parallel: A guide to Efficient Multi-GPU Training

Read the original on Hugging Face Blog →

The Flow has not summarised this story yet — read it at Hugging Face Blog.

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
Jun 2

SUPREME: A Multi-GPU Framework for Reproducible Image Unlearning Method Evaluation

arXiv:2606. 00380v1 Announce Type: cross Abstract: Machine unlearning removes the influence of specific training data from a trained model without retraining it from scratch.

By Petros Andreou, Jamie Lanyon, Axel Finke, Georgina Cosma
computer-vision
More like this →
arXiv Machine Learning
Aug 6

SparseDitto: Customizing GPU Kernels for Different Sparsity Patterns with LLM-Based Agentic System

arXiv:2608. 05033v1 Announce Type: cross Abstract: Sparse matrix kernels are fundamental to scientific computing, graph analytics, and machine learning.

By Shiyang Li, Guangyan Sun, Jinwei Tang, Yanzhi Wang, Mingyi Hong, Caiwen Ding
llmsagentsefficiency
More like this →
arXiv Machine Learning
Jun 30

GPU Parallelization Strategies for Forward and Backward Propagation in Shallow Neural Networks: A CUDA-Based Comparative Study

arXiv:2606. 30497v1 Announce Type: cross Abstract: We present a comparative study of CUDA optimization strategies applied to forward and backward propagation in a shallow neural network.

By Rania Zitouni, Nadine Bousdjira, Sarah Hasnaoui, Amel Sadoun, Fatma Salhi
More like this →
arXiv Machine Learning
Jun 8

TorchKM: A GPU-Oriented Library for Kernel Learning and Model Selection

arXiv:2606. 06742v1 Announce Type: new Abstract: TorchKM is an open-source library for kernel machines, including support vector machines, kernel logistic regression, and kernel quantile regression, with GPU acceleration.

By Yikai Zhang, Gaoxiang Jia, Jie Ding, Boxiang Wang
benchmarks
More like this →
arXiv AI
Jul 3

Mixture-of-Parallelisms: Towards Memory-Efficient Training Stack for Mixture-of-Experts Models

arXiv:2607. 01844v1 Announce Type: cross Abstract: This paper showcases a memory-efficient training stack for Mixture-of-Experts (MoE) models.

By Xuan-Phi Nguyen, Shrey Pandit, Yiran Zhao, Semih Yavuz, Silvio Savarese, Shafiq Joty
fine-tuningefficiencybenchmarks
More like this →
arXiv AI
Jul 1

An Efficient Heterogeneous Co-Design for Fine-Tuning on a Single GPU

arXiv:2603. 16428v2 Announce Type: replace-cross Abstract: Fine-tuning Large Language Models (LLMs) has become essential for domain adaptation, but its memory-intensive property exceeds the capabilities of most GPUs.

By Ruijia Yang, Zeyi Wen
llmsfine-tuningefficiency
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.0.0 · bb4ee0e