Incredibly Fast BLOOM Inference with DeepSpeed and Accelerate
Related stories
Fast Inference on Large Language Models: BLOOMZ on Habana Gaudi2 Accelerator
Making thousands of open LLMs bloom in the Vertex AI Model Garden
Atompack: A Storage and Distribution Layer for Read-Heavy Atomistic ML Training Datasets
arXiv:2606. 29975v1 Announce Type: new Abstract: Atomistic machine learning datasets are increasingly used for training: large immutable snapshots are read repeatedly, shuffled across epochs, staged across clusters' storage systems, and republished as reusable scientific artifacts.
scDataset: Scalable Data Loading for Deep Learning on Large-Scale Single-Cell Omics
arXiv:2506. 01883v3 Announce Type: replace-cross Abstract: Training deep learning models on single-cell datasets with hundreds of millions of cells requires loading data from disk, as these datasets exceed available memory.
Fast BIB simulation at a future Muon Collider with generative machine learning
The paper presents the first machine learning models for fast generation of beam‑induced background (BIB) in tracking detectors at a future Muon Collider. Two architectures are explored: a high‑fidelity tabular diffusion model and a faster circular spline flow model. Both produce BIB hits and tracks that closely match full simulation results, achieving over an order of magnitude speed‑up while requiring far less computational resources.
Atompack: A Storage and Distribution Layer for Read-Heavy Atomistic ML Training Datasets
Atomistic machine learning datasets are increasingly used for training: large immutable snapshots are read repeatedly, shuffled across epochs, staged across clusters' storage systems, and republished as reusable scientific artifacts. This workload differs from interactive scientific curation, where mutable records and ad hoc inspection are often more important than random indexed throughput.
Up to 3.2x Faster Inference with LFM2.5-DSpark
DARTree: Speculative Diffusion Decoding with Autoregressive Draft Trees
Speculative decoding losslessly accelerates autoregressive language models by verifying multiple draft tokens in parallel. Diffusion-based drafters further reduce proposal latency by predicting an entire token block in parallel, but their position-wise distributions are marginal rather than conditioned on tokens selected along each draft path.
Dion3: Full-Stack Orthogonal Updates
arXiv:2608. 11612v1 Announce Type: cross Abstract: The Muon optimizer incurs a significant overhead cost due to its cubic-time Newton-Schulz orthogonalization step.
Adapting Tree-Structured Speculative Decoding to DeepSeek-V4 for Efficient Inference
arXiv:2609.24698v1 Announce Type: new Abstract: Repeated execution of the target model during autoregressive decoding is a major source of LLM inference latency. Unlike linear speculation, which foll...
Efficient Iterative Retrieval with Heterogeneous Batching
Orthrus is a serving system that performs heterogeneous batching of embedding and generative models within a single inference loop. It uses chunked embedding with incremental pooling and workload‑aware batch composition to unify conflicting computational patterns. Experiments on four A100 GPUs show that Orthrus improves throughput by 1.28×–4.52× and reduces p99 latency by up to 55.8% compared to baseline deployments.