arXiv Machine Learning
Aug 28

COFM: Consistent Optimal Transport Flow Matching via Partially Input Convex Neural Networks

The paper introduces COFM, a framework for consistent optimal transport flow matching that uses partially input convex neural networks (PICNN) to parameterize the transport potential. By adding a Hamilton‑Jacobi residual to the training objective, COFM enforces dynamical consistency and supports both one‑step transport and multi‑step ODE sampling without costly inner optimization. Experiments on benchmark datasets show that COFM achieves competitive performance while reducing L^2‑UVP by over 2× and cutting computational time by about 9× compared to state‑of‑the‑art models.

By Fanghui Song, Zhongjian Wang, Jiebao Sun
arXiv Machine Learning
Aug 19

Composing Flow-Matching Energies with Known Physics: Generation, OOD Detection, and Inversion on PDE Fields

The paper presents a method that combines flow‑matching models with energy‑based modeling to explicitly construct scalar energy functions for physical fields. These energies are derived from a matching regression objective on a linear Gaussian interpolation, avoiding variational approximations or extra MCMC steps, and can be used for energy‑corrected data generation, out‑of‑distribution detection, and posterior sampling in inverse problems. The approach enables general MCMC samplers that reduce PDE residuals and spectral distance, and it demonstrates that combining data‑driven and physics‑based energies improves OOD detection accuracy.

By Yixuan Sun, Anirban Samaddar, Sandeep Madireddy
arXiv Machine Learning
Aug 26

Generalization, memorization, and overfitting for diffusion models trained in the lazy high-dimensional regime

The paper investigates diffusion models trained in a lazy high‑dimensional regime, extending benign overfitting theory to generative settings. By analyzing denoising score matching in a vector‑valued RKHS with an inner‑product kernel, the authors derive exact risk trajectories under gradient flow when the number of samples scales proportionally with dimensionality. These trajectories reveal three distinct phases—spectral generalization, noise‑dominated interpolation, and empirical Bayes memorization—whose interplay shapes the distribution of generated samples.

By Hugo Latourelle-Vigeant, Sinho Chewi, Aram-Alexandre Pooladian, John Sous, Theodor Misiakiewicz