arXiv Machine Learning

Multivariate Scientific Data Compression with Learned Cross-Variable Latent Decorrelation and Autoregressive Entropy Modeling

arXiv Machine Learning
Sep 22

Neural Residual Modeling for Scientific Data Compression under Guaranteed Error Bounds

The paper introduces a new compression pipeline for scientific simulation data that combines Residual Vector Quantization (RVQ) with a U‑Net post‑processing network to correct pixel‑space residuals, followed by a Guaranteed Autoencoder (GAE) that enforces block‑wise error bounds. The U‑Net is trained to predict spatially structured residuals, addressing limitations of latent‑space only approaches. Experiments on S3D, JHTDB, and E3SM datasets show improved NRMSE and compression ratios compared to RVQ alone while maintaining strict error guarantees.

By Surya Majumder, Liangji Zhu, Sanjay Ranka, Anand Rangarajan
arXiv Machine Learning
4d ago

HALO: Enhancing Time Series Generation via Hyperspherical Latents and Masked AutoregRessive Modeling

HALO introduces a hyperspherical VAE to constrain continuous latent representations to a fixed‑radius shell, stabilizing numerical fluctuations. It then employs a masked autoregressive model that balances parallel decoding with temporal correlation learning, reducing inference steps and improving stability. Experiments show HALO achieves state‑of‑the‑art generation performance with significantly better inference efficiency compared to existing baselines.

By Chunyi Hou, Xiangfei Qiu, Hanyin Cheng, Yutong Li, Bin Yang
arXiv Machine Learning
Sep 25

Beyond Compression: Training Latent Representations for Stable Long-Horizon Rollout in Neural Surrogate Solvers

The paper investigates why latent neural surrogate solvers, which compress physical system dynamics into a lower‑dimensional space, often fail during long‑horizon autoregressive rollouts. It demonstrates that training the latent representation only for reconstruction leads to instability, and proposes a set of training interventions—Koopman operator learning, Hamming noise injection, and multi‑step rollout fine‑tuning—that align the latent space with long‑horizon forecasting. These interventions reduce long‑rollout error by about 40 % and achieve accuracy comparable to full‑resolution models while using far fewer floating‑point operations and GPU memory, enabling stable extrapolation in mesoscale crystal‑plasticity simulations of high‑cycle fatigue.

By Andreas E. Robertson, Ashley T. Lenau, John D. Shimanek, Benjamin A. Jasperson, Vivek Oommen, David L. Damm, Krishna Garikipati, Remi Dingreville
arXiv Machine Learning
Aug 28

SimCast-S2S: An Efficient Generative Model for Subseasonal Precipitation Forecasting via Transfer Learning from Climate Simulations

SimCast‑S2S is a generative latent‑diffusion model designed for probabilistic subseasonal‑to‑seasonal precipitation forecasting. It tackles three key challenges: it uses a diffusion pipeline to capture uncertainty, operates in a compact latent space to enable efficient large‑ensemble generation, and leverages transfer learning with low‑rank adaptation to train on limited reanalysis data after pretraining on climate simulations. The model outperforms deep‑learning baselines and competes with, or surpasses, operational systems such as the ECMWF‑S2S baseline without requiring extensive post‑processing.

By Hiep V. Dang, Antonios Mamalakis
arXiv Machine Learning
Jun 10

When to Align, When to Predict: A Phase Diagram for Multimodal Learning

arXiv:2606. 11190v1 Announce Type: new Abstract: Cross-modal alignment (CA) and cross-modal prediction (CP) are the dominant paradigms for multimodal representation learning, yet there is no systematic understanding of when each succeeds, when each fails, and when cross-modal training helps at all -- a gap that leaves practitioners, especially in scientific domains like biomedicine or astrophysics, with heterogeneous instruments and multiple levels of organization and measurement, unable to diagnose why standard methods underperform the best single modality.

By Ilay Kamai, Hugues Van Assel, Aviv Regev, Hagai B. Perets, Randall Balestriero