Training Variable Long Sequences with Data-Centric Parallel
arXiv:2608. 07524v1 Announce Type: new Abstract: Training deep learning models on variable long sequences poses significant computational challenges.
arXiv:2506. 01883v3 Announce Type: replace-cross Abstract: Training deep learning models on single-cell datasets with hundreds of millions of cells requires loading data from disk, as these datasets exceed available memory.
arXiv:2608. 07524v1 Announce Type: new Abstract: Training deep learning models on variable long sequences poses significant computational challenges.
Orthrus is a serving system that performs heterogeneous batching of embedding and generative models within a single inference loop. It uses chunked embedding with incremental pooling and workload‑aware batch composition to unify conflicting computational patterns. Experiments on four A100 GPUs show that Orthrus improves throughput by 1.28×–4.52× and reduces p99 latency by up to 55.8% compared to baseline deployments.
The paper introduces NIO Bench, a benchmarking framework that profiles storage I/O for six machine learning model types, using Python hooks and Linux strace to capture detailed access patterns. Experiments on a Ceph-backed Kubernetes cluster show that I/O is dominated by data preparation, model loading, and checkpointing, with training becoming compute-bound once data is staged. The study finds a power‑law distribution of file usage and identifies cache‑miss read tail latency as the main storage bottleneck, recommending aggressive prefetching, page‑cache pinning, and bursty write handling for ML‑optimized storage.
arXiv:2606. 29975v1 Announce Type: new Abstract: Atomistic machine learning datasets are increasingly used for training: large immutable snapshots are read repeatedly, shuffled across epochs, staged across clusters' storage systems, and republished as reusable scientific artifacts.
Atomistic machine learning datasets are increasingly used for training: large immutable snapshots are read repeatedly, shuffled across epochs, staged across clusters' storage systems, and republished as reusable scientific artifacts. This workload differs from interactive scientific curation, where mutable records and ad hoc inspection are often more important than random indexed throughput.
DeepSeek‑V4.1‑Flash is a multimodal Mixture‑of‑Experts model with 552 B backbone parameters that supports contexts of up to one million tokens. It uses a Causal Encoder‑Decoder architecture that activates 16 B parameters per token during decoding but only 8 B during prefill, improving cost efficiency for agentic workloads. The model introduces cross‑layer KV cache reuse via Compressed Sparse Attention 2 and FP4 KV caching, reducing its global KV cache footprint to 890 bytes per token and its persistent footprint to roughly one‑eighth of the previous version, while still delivering superior performance after pretraining on a 45 T‑token multimodal corpus.
The paper proposes a Learnware-based framework for deploying scene‑specific CSI feedback models in 6G systems. A centralized AI data center maintains a catalog of pre‑trained models, each tagged with semantic and statistical specifications. Base stations retrieve the most relevant model using only statistical fingerprints, which reduces data privacy risks, lowers retrieval latency, and cuts fine‑tuning effort, achieving up to 57.7% performance gains over a general model.
arXiv:2607. 18187v1 Announce Type: cross Abstract: Large-scale scientific simulations generate volumetric data at rates that far outpace advances in storage and network bandwidth, making effective lossy compression increasingly critical.
arXiv:2604. 24806v2 Announce Type: replace-cross Abstract: Modern Deep Learning Recommendation Models (DLRMs) follow scaling laws with sequence length, driving the frontier toward ultra-long User Interaction History (UIH).
arXiv:2607. 15745v1 Announce Type: new Abstract: Common practice when training Convolutional Neural Networks (CNNs) is to use randomly shuffled mini-batches.
arXiv:2606. 01155v1 Announce Type: cross Abstract: Scaling laws for dense LLMs under infinite data are well explored, but how sparsity interacts with limited data is not.
LeanStream introduces a speculate‑and‑refine streaming framework that enables efficient on‑device inference of large language models by progressively refining computation, loading, and cache‑retention priorities using partial GPU results. This approach allows fine‑grained overlap between GPU execution and storage I/O, avoiding the trade‑offs of existing systems that serialize execution or incur redundant weight fetches. Implemented on mobile and embedded platforms, LeanStream reduces memory usage by 4.8× to 7.5× compared to prior work while improving token generation throughput by 1.6× to 2.1×.