Hugging Face Blog

Incredibly Fast BLOOM Inference with DeepSpeed and Accelerate

arXiv Machine Learning
Jun 30

Atompack: A Storage and Distribution Layer for Read-Heavy Atomistic ML Training Datasets

arXiv:2606. 29975v1 Announce Type: new Abstract: Atomistic machine learning datasets are increasingly used for training: large immutable snapshots are read repeatedly, shuffled across epochs, staged across clusters' storage systems, and republished as reusable scientific artifacts.

By Ali Ramlaoui, Daniel T. Speckhard, Sagar Pal, Fragkiskos D. Malliaros, Alexandre Duval, Victor Schmidt
arXiv Machine Learning
Sep 14

Fast BIB simulation at a future Muon Collider with generative machine learning

The paper presents the first machine learning models for fast generation of beam‑induced background (BIB) in tracking detectors at a future Muon Collider. Two architectures are explored: a high‑fidelity tabular diffusion model and a faster circular spline flow model. Both produce BIB hits and tracks that closely match full simulation results, achieving over an order of magnitude speed‑up while requiring far less computational resources.

By Radha Mastandrea, Shiyu Peng, Benjamin Rosser, Matt LeBlanc
Hugging Face Trending Papers
Jun 29

Atompack: A Storage and Distribution Layer for Read-Heavy Atomistic ML Training Datasets

Atomistic machine learning datasets are increasingly used for training: large immutable snapshots are read repeatedly, shuffled across epochs, staged across clusters' storage systems, and republished as reusable scientific artifacts. This workload differs from interactive scientific curation, where mutable records and ad hoc inspection are often more important than random indexed throughput.

Hugging Face Trending Papers
Aug 13

DARTree: Speculative Diffusion Decoding with Autoregressive Draft Trees

Speculative decoding losslessly accelerates autoregressive language models by verifying multiple draft tokens in parallel. Diffusion-based drafters further reduce proposal latency by predicting an entire token block in parallel, but their position-wise distributions are marginal rather than conditioned on tokens selected along each draft path.

arXiv AI
Aug 13

Dion3: Full-Stack Orthogonal Updates

arXiv:2608. 11612v1 Announce Type: cross Abstract: The Muon optimizer incurs a significant overhead cost due to its cubic-time Newton-Schulz orthogonalization step.

By Noah Amsel, Jack Zhang, Kwangjun Ahn, Ali Naeimi, Austin Feng, Berlin Chen, Tri Dao, John Langford
arXiv Computation and Language
Sep 23

Efficient Iterative Retrieval with Heterogeneous Batching

Orthrus is a serving system that performs heterogeneous batching of embedding and generative models within a single inference loop. It uses chunked embedding with incremental pooling and workload‑aware batch composition to unify conflicting computational patterns. Experiments on four A100 GPUs show that Orthrus improves throughput by 1.28×–4.52× and reduces p99 latency by up to 55.8% compared to baseline deployments.

By Dohyun Park, Hubertus Franke, Daniel G. Waddington, Swaminathan Sundararaman, Yongjoo Park