Residual Modeling for High-Fidelity Learned Compression of Scientific Data
arXiv:2606. 05389v1 Announce Type: new Abstract: Lossy compression is essential for massive spatiotemporal data from scientific simulations.
arXiv:2607. 18187v1 Announce Type: cross Abstract: Large-scale scientific simulations generate volumetric data at rates that far outpace advances in storage and network bandwidth, making effective lossy compression increasingly critical.
arXiv:2606. 05389v1 Announce Type: new Abstract: Lossy compression is essential for massive spatiotemporal data from scientific simulations.
arXiv:2606. 14353v1 Announce Type: new Abstract: Error-bounded lossy compression is a fundamental technique for managing the rapidly growing volumes of scientific data produced by modern simulations and observational instruments.
arXiv:2607. 01164v1 Announce Type: new Abstract: Recent work has shown that implicit neural representations (INRs) can be trained to effectively compress structured and unstructured volume data, allowing for direct data querying with a reduced memory footprint.
arXiv:2606. 15959v1 Announce Type: cross Abstract: Neural networks are used as generative surrogate models for scientific discovery, which are trainable approximations of scientific simulations.
arXiv:2608. 14112v1 Announce Type: cross Abstract: Scientific simulations often produce scalar volumes faster than they can be stored, transferred, and loaded, while in situ reduction must use only a limited share of simulation resources.
arXiv:2607. 11883v1 Announce Type: new Abstract: Compression is fundamental to intelligence.
arXiv:2605. 17985v2 Announce Type: replace-cross Abstract: We propose a new method for compressing physics foundation models (PFMs) which is a new trend in AI for Science.
arXiv:2606. 29975v1 Announce Type: new Abstract: Atomistic machine learning datasets are increasingly used for training: large immutable snapshots are read repeatedly, shuffled across epochs, staged across clusters' storage systems, and republished as reusable scientific artifacts.
arXiv:2606. 09659v1 Announce Type: cross Abstract: Long-context language model inference is bottlenecked by memory, as the KV cache grows with context length.
Atomistic machine learning datasets are increasingly used for training: large immutable snapshots are read repeatedly, shuffled across epochs, staged across clusters' storage systems, and republished as reusable scientific artifacts. This workload differs from interactive scientific curation, where mutable records and ad hoc inspection are often more important than random indexed throughput.
Compression is fundamental to intelligence. A model that can represent its training data as a short code has discovered regularities that enable generalization.
arXiv:2510. 20535v2 Announce Type: replace-cross Abstract: Recent techniques such as retrieval-augmented generation or chain-of-thought reasoning have led to longer contexts and increased inference costs.