Residual Modeling for High-Fidelity Learned Compression of Scientific Data
arXiv:2606. 05389v1 Announce Type: new Abstract: Lossy compression is essential for massive spatiotemporal data from scientific simulations.
arXiv:2606. 05389v1 Announce Type: new Abstract: Lossy compression is essential for massive spatiotemporal data from scientific simulations.
The paper introduces a new compression pipeline for scientific simulation data that combines Residual Vector Quantization (RVQ) with a U‑Net post‑processing network to correct pixel‑space residuals, followed by a Guaranteed Autoencoder (GAE) that enforces block‑wise error bounds. The U‑Net is trained to predict spatially structured residuals, addressing limitations of latent‑space only approaches. Experiments on S3D, JHTDB, and E3SM datasets show improved NRMSE and compression ratios compared to RVQ alone while maintaining strict error guarantees.
arXiv:2606. 14353v1 Announce Type: new Abstract: Error-bounded lossy compression is a fundamental technique for managing the rapidly growing volumes of scientific data produced by modern simulations and observational instruments.
arXiv:2607. 18187v1 Announce Type: cross Abstract: Large-scale scientific simulations generate volumetric data at rates that far outpace advances in storage and network bandwidth, making effective lossy compression increasingly critical.
arXiv:2606. 03279v1 Announce Type: new Abstract: In AI for Science, physics-informed losses are increasingly used to train learned compressors for scientific data, but their rate-distortion implications remain poorly understood.
HALO introduces a hyperspherical VAE to constrain continuous latent representations to a fixed‑radius shell, stabilizing numerical fluctuations. It then employs a masked autoregressive model that balances parallel decoding with temporal correlation learning, reducing inference steps and improving stability. Experiments show HALO achieves state‑of‑the‑art generation performance with significantly better inference efficiency compared to existing baselines.
arXiv:2606. 06576v1 Announce Type: new Abstract: In the sciences, regression tasks often require predicting high-dimensional outputs from few training examples.
arXiv:2606. 11691v1 Announce Type: new Abstract: Latent diffusion and flow matching have emerged as leading approaches for synthetic turbulence generation, yet they systematically under-represent dissipation-range amplitudes.
The paper investigates why latent neural surrogate solvers, which compress physical system dynamics into a lower‑dimensional space, often fail during long‑horizon autoregressive rollouts. It demonstrates that training the latent representation only for reconstruction leads to instability, and proposes a set of training interventions—Koopman operator learning, Hamming noise injection, and multi‑step rollout fine‑tuning—that align the latent space with long‑horizon forecasting. These interventions reduce long‑rollout error by about 40 % and achieve accuracy comparable to full‑resolution models while using far fewer floating‑point operations and GPU memory, enabling stable extrapolation in mesoscale crystal‑plasticity simulations of high‑cycle fatigue.
SimCast‑S2S is a generative latent‑diffusion model designed for probabilistic subseasonal‑to‑seasonal precipitation forecasting. It tackles three key challenges: it uses a diffusion pipeline to capture uncertainty, operates in a compact latent space to enable efficient large‑ensemble generation, and leverages transfer learning with low‑rank adaptation to train on limited reanalysis data after pretraining on climate simulations. The model outperforms deep‑learning baselines and competes with, or surpasses, operational systems such as the ECMWF‑S2S baseline without requiring extensive post‑processing.
arXiv:2608. 04222v1 Announce Type: cross Abstract: Turbulence is a central testbed for machine learning on physical dynamics because its governing laws are known exactly.
arXiv:2606. 11190v1 Announce Type: new Abstract: Cross-modal alignment (CA) and cross-modal prediction (CP) are the dominant paradigms for multimodal representation learning, yet there is no systematic understanding of when each succeeds, when each fails, and when cross-modal training helps at all -- a gap that leaves practitioners, especially in scientific domains like biomedicine or astrophysics, with heterogeneous instruments and multiple levels of organization and measurement, unable to diagnose why standard methods underperform the best single modality.