From Files to Chunks: Improving HF Storage Efficiency
Related stories
From Chunks to Blocks: Accelerating Uploads and Downloads on the Hub
Smart Compaction: Predicting Compaction Utility from Lakehouse Table Metadata
arXiv:2608. 08639v1 Announce Type: new Abstract: Open lakehouse table formats accumulate small data files over time, which degrades query performance.
LayoutBench: Performance Benchmarking of Cloud Storage Layouts for Multimedia Data
arXiv:2607. 28880v1 Announce Type: cross Abstract: Modern multimedia machine learning workloads increasingly store large-scale datasets in cloud object storage services such as AWS S3.
Atompack: A Storage and Distribution Layer for Read-Heavy Atomistic ML Training Datasets
Atomistic machine learning datasets are increasingly used for training: large immutable snapshots are read repeatedly, shuffled across epochs, staged across clusters' storage systems, and republished as reusable scientific artifacts. This workload differs from interactive scientific curation, where mutable records and ad hoc inspection are often more important than random indexed throughput.
Atompack: A Storage and Distribution Layer for Read-Heavy Atomistic ML Training Datasets
arXiv:2606. 29975v1 Announce Type: new Abstract: Atomistic machine learning datasets are increasingly used for training: large immutable snapshots are read repeatedly, shuffled across epochs, staged across clusters' storage systems, and republished as reusable scientific artifacts.
RedKnot-MLA: Multi-Head Offline-Online Reuse for DeepSeek-V4 Long-Context Serving
arXiv:2609.07008v1 Announce Type: new Abstract: Multi-head latent attention (MLA) exposes many logical query heads through one packed latent KV stream. This representation is memory efficient, but it...
Breaking Database Lock-in: Agentic Regeneration of High Performance Storage Readers for Database Bypass
arXiv:2607. 07696v1 Announce Type: cross Abstract: Analytical workloads operating on data stored in external database systems face a fundamental bottleneck: data access is guarded entirely by the database driver, like JDBC or ODBC, forcing all reads through query execution and other driver layers that are not designed for bulk columnar analytics.
Document Retrieval-Aware Chunking (D-RAC): Universal Retrieval-Aware Ingestion of Enterprise Documents via PDF Normalization and Multimodal Markdown Conversion
arXiv:2609.24220v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) systems over enterprise knowledge bases must ingest heterogeneous document formats -- PDFs, Word documents, presen...
Breaking Database Lock-in: Agentic Regeneration of High Performance Storage Readers for Database Bypass
Analytical workloads operating on data stored in external database systems face a fundamental bottleneck: data access is guarded entirely by the database driver, like JDBC or ODBC, forcing all reads through query execution and other driver layers that are not designed for bulk columnar analytics. We present Jailbreak, an approach that bypasses the database engine entirely by reading storage files directly and materializing data as in-memory columnar buffers.
Benchmarking Storage Systems for Machine Learning Workloads Using NIO Bench
The paper introduces NIO Bench, a benchmarking framework that profiles storage I/O for six machine learning model types, using Python hooks and Linux strace to capture detailed access patterns. Experiments on a Ceph-backed Kubernetes cluster show that I/O is dominated by data preparation, model loading, and checkpointing, with training becoming compute-bound once data is staged. The study finds a power‑law distribution of file usage and identifies cache‑miss read tail latency as the main storage bottleneck, recommending aggressive prefetching, page‑cache pinning, and bursty write handling for ML‑optimized storage.
Characterizing High Bandwidth Flash for LLM Serving
arXiv:2609.39131v1 Announce Type: new Abstract: Large language model (LLM) serving requires substantial memory to store model weights and KV caches. As models grow larger and contexts become longer,...