arXiv Machine Learning By Ali Ramlaoui, Daniel T. Speckhard, Sagar Pal, Fragkiskos D. Malliaros, Alexandre Duval, Victor Schmidt

Atompack: A Storage and Distribution Layer for Read-Heavy Atomistic ML Training Datasets

Read the original on arXiv Machine Learning →

arXiv:2606. 29975v1 Announce Type: new Abstract: Atomistic machine learning datasets are increasingly used for training: large immutable snapshots are read repeatedly, shuffled across epochs, staged across clusters' storage systems, and republished as reusable scientific artifacts.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
Jun 29

Atompack: A Storage and Distribution Layer for Read-Heavy Atomistic ML Training Datasets

Atomistic machine learning datasets are increasingly used for training: large immutable snapshots are read repeatedly, shuffled across epochs, staged across clusters' storage systems, and republished as reusable scientific artifacts. This workload differs from interactive scientific curation, where mutable records and ad hoc inspection are often more important than random indexed throughput.

arXiv Machine Learning
Sep 10

Benchmarking Storage Systems for Machine Learning Workloads Using NIO Bench

The paper introduces NIO Bench, a benchmarking framework that profiles storage I/O for six machine learning model types, using Python hooks and Linux strace to capture detailed access patterns. Experiments on a Ceph-backed Kubernetes cluster show that I/O is dominated by data preparation, model loading, and checkpointing, with training becoming compute-bound once data is staged. The study finds a power‑law distribution of file usage and identifies cache‑miss read tail latency as the main storage bottleneck, recommending aggressive prefetching, page‑cache pinning, and bursty write handling for ML‑optimized storage.

By Jonathan W. Morris, Ionut Mistreanu, Connor Louie