arXiv:2506. 01883v3 Announce Type: replace-cross Abstract: Training deep learning models on single-cell datasets with hundreds of millions of cells requires loading data from disk, as these datasets exceed available memory.
By Davide D'Ascenzo, Sebastiano Cultrera di Montesano
arXiv:2505. 06835v5 Announce Type: replace Abstract: Sliced optimal transport (SOT), or sliced Wasserstein (SW) distance, is widely recognized for its statistical and computational scalability.
By Khai Nguyen
arXiv:2607. 28880v1 Announce Type: cross Abstract: Modern multimedia machine learning workloads increasingly store large-scale datasets in cloud object storage services such as AWS S3.
By Debopam Sanyal, Hongjie Chen, Alexey Tumanov, Joshua Kimball
arXiv:2608. 00720v1 Announce Type: cross Abstract: Mapping neural networks to FPGAs enables low-latency, energy-efficient inference, particularly for lookup table (LUT)-based models that eliminate multipliers and map directly to reconfigurable fabric.
By Oliver Cassidy, Marta Andronic, George A. Constantinides
arXiv:2502. 06434v2 Announce Type: replace-cross Abstract: Dataset pruning (DP) and dataset distillation (DD) fundamentally differ in their outputs: DP selects original image subsets, while DD generates synthetic images.
By Lingao Xiao, Songhua Liu, Yang He, Xinchao Wang
arXiv:2601. 07048v5 Announce Type: replace-cross Abstract: Approximate nearest neighbor search (ANNS) is a core problem in machine learning and information retrieval applications.
By Hunter McCoy, Zikun Wang, Prashant Pandey
arXiv:2606. 18209v1 Announce Type: new Abstract: Dataset distillation (DD) has emerged as a prominent approach in data centric machine learning, aiming to synthesize compact training sets for efficient training by compressing the information in large datasets into a small number of synthetic samples.
By Trisha Mittal, Akshay Mehra, Joshua Kimball
arXiv:2606. 02142v1 Announce Type: new Abstract: The ongoing digitization has led to a proliferation of time-series data streams that monitor a variety of processes, from which valuable insights may be obtained.
By David Campos, Bin Yang, Tung Kieu, Lei Chen, Chenjuan Guo, Christian S. Jensen
arXiv:2606. 15346v1 Announce Type: cross Abstract: Spatio-temporal prediction supports radar/satellite nowcasting and city-scale traffic monitoring, but modern models are often too expensive for real-time deployment.
By Fuyan Zhang, Yuqi Li, Yingli Tian, Edmond S. L. Ho
arXiv:2602. 22101v3 Announce Type: replace-cross Abstract: Many real-world applications generate continuous data streams for regression.
By Pantia-Marina Alchirch, Dimitrios I. Diochnos
arXiv:2606. 11761v1 Announce Type: new Abstract: Dynamic data pruning techniques aim to reduce computational cost while minimizing information loss by periodically selecting representative subsets of input data during model training.
By Atif Hassan, Swanand Khare, Jiaul H. Paik