arXiv Machine Learning

Time series generation with spectrally aligned latent flow matching

The paper introduces a spectrally-aligned latent-flow model for time‑series generation that trains the latent space to preserve dynamical properties relevant to synthetic data quality. By incorporating fine‑tuning losses based on Fourier, wavelet, and signature transforms, the method mitigates spectral mismatches caused by latent compression and ensures alignment with true signals in terms of smoothness and targeted spectral content. Experiments on real‑world long‑range univariate and multivariate benchmarks show that the aligned model outperforms a base latent‑flow model and state‑of‑the‑art approaches in signal realism, computational efficiency, and local structure alignment.

arXiv AI
1d ago

Wavelet Flow Matching for Time Series

The paper introduces Wavelet Flow Matching, a method for generating multivariate time series by applying flow matching to multilevel discrete wavelet coefficients. By working in the wavelet domain, the model captures coarse-to-fine temporal structure implicitly and uses a channel-token transformer to model cross-channel dependencies. Experiments on seven benchmark datasets and four sequence lengths show that the approach matches or surpasses existing methods, especially in Context-FID and discriminative score metrics.

By Lucas Poinsignon, Jorge da Silva Gon\c{c}alves, Samuel Ruip\'erez-Campillo, Julia E. Vogt
arXiv Machine Learning
Sep 24

On the Diffusibility of High-Dimensional Latents

The paper investigates how fine‑tuning pretrained visual encoders for faithful image reconstruction affects diffusion models that operate in the resulting latent space. It finds that such fine‑tuning reduces the effective dimensionality of the latent representation, causing standard velocity‑prediction flow‑matching to fit noise outside the low‑dimensional signal manifold and making optimization inefficient. Consequently, the authors propose using a clean‑data ($oldsymbol{x}_{0}$) parameterization, which focuses learning on the signal manifold and consistently improves text‑to‑image generation across multiple strong‑reconstruction encoders.

By Chao Feng, Zhiyang Xu, Bowei Chen, Yuanjun Xiong, Xiyao Wang, Jui-Hsien Wang, Richard Zhang, Zhe Lin, Andrew Owens, Yijun Li
arXiv Computer Vision
Sep 3

Balancing Frequencies and Pixels in Flow Matching

The paper introduces a Focal Log-Frequency Loss (f-loss) to counteract the spectral imbalance in pixel-space flow matching, where low frequencies dominate training. By balancing learning signals across frequencies and combining early frequency-domain supervision with later pixel-space refinement, the method accelerates convergence by up to 40% and improves FID and perceptual fidelity across multiple model scales. It requires no architectural changes and can replace existing flow matching losses as a drop‑in solution.

By Lucas Degeorge, Paul Couairon, Arijit Ghosh, Alexei A. Efros, David Picard, Vicky Kalogeiton
arXiv AI
Sep 16

SOTER: A Generative Time-Series Foundation Model for Wearable Human Physiological Signals

SOTER is a generative foundation model designed for wearable physiological time‑series data. It integrates cross‑channel coupling, spectrum‑guided expert specialization, and continuous‑time latent evolution, using a spatial feature‑aware backbone, a PSD‑guided mixture‑of‑experts layer, and a neural controlled differential equation decoder. Trained on 226 billion time points from five public datasets, SOTER outperforms baselines in zero‑shot forecasting, classification, and imputation across six benchmarks, and remains robust to additive noise.

By Fangke Chen, Sirry Chen, Wei Chen, Zhongyu Wei
arXiv Computer Vision
Sep 18

Training Flow Matching: The Role of Weighting and Parameterization

The paper investigates training objectives for denoising-based generative models, focusing on loss weighting and output parameterization such as noise-, clean image-, and velocity-based formulations. It conducts a systematic numerical study across synthetic datasets with controlled geometry and real image data, evaluating denoising accuracy via PSNR and generative quality via FID. The goal is to disentangle how training choices interact with data manifold dimensionality, model architecture, and dataset size, offering practical design insights rather than proposing a new method.

By Anne Gagneux, S\'egol\`ene Martin, R\'emi Gribonval, Mathurin Massias