arXiv Machine Learning By Silas Koemen

Conditioning Tree-Based Diffusions and Flows for Probabilistic Tabular Regression

Read the original on arXiv Machine Learning →

arXiv:2607. 28864v1 Announce Type: cross Abstract: Tree-based diffusion models fit flexible conditional predictive distributions for tabular regression without a neural density estimator, but they inherit their design defaults---noising path, parameterization, training distribution, features, sampler---from the neural setting.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computer Vision
Aug 24

Difficulty-Calibrated Interpolation Paths for Conditional Flow Matching

The paper introduces Difficulty-Calibrated Flow Matching, a method that adapts the noise-to-data interpolation schedule in Conditional Flow Matching based on a pilot run’s loss profile. By setting the schedule to the quantile function of this difficulty profile, the training trajectory spends more time where the velocity is hardest to learn. Experiments on CIFAR-10, MNIST, and Fashion‑MNIST show that this calibrated path achieves the best FID on CIFAR‑10 and outperforms all fixed schedules in large‑batch, few‑update settings, where compute is most limited.

By Airin Akter Tania, Md Raihan Khan
arXiv Machine Learning
Sep 22

Optimizers for Diffusion Models: A Controlled Benchmark

The paper introduces the first controlled benchmark for optimizers in discrete diffusion models, evaluating seven optimizers (AdamW, Lion, Muon, SOAP, MARS, MARS‑M, Schedule‑Free) across four diffusion formulations: masked diffusion on text8, uniform diffusion on QM9 and LM1B, and Gaussian diffusion on CelebA‑64. Each optimizer undergoes the same search protocol and is retrained with full budget and multiple seeds, revealing that AdamW, while strong, is not universally optimal and that optimizers validated on autoregressive language models (Muon, MARS‑M, SOAP) can outperform tuned AdamW on certain tasks.

By Arman Bolatov, Egor Shulgin, David Li, Abduragim Shtanchaev, Sebastian U. Stich, Maxim Panov, Eric Moulines, Peter Richt\'arik, Martin Tak\'a\v{c}
arXiv Machine Learning
Aug 5

Any-OPD: Heterogeneous On-Policy Distillation for Flow-Matching Models via Representation-Space Bridging

arXiv:2608. 03316v1 Announce Type: new Abstract: On-policy distillation, in which a teacher corrects samples that the student itself generates, presupposes that the two models speak the same language: identical VAE latents, matching architectures, and a common timestep grid.

By Siming Fu, Zheming Fu, Ruizhe He, Hualiang Wang, Jie Huang, Xiaoxiao Ma, Mingchen Zhong, Weihu Huang, Xiaoxuan He, Haojun Xu
arXiv AI
Sep 3

GeoSPRINT: Geometric Redundancy-Aware Step Pruning for Inference in Diffusion Trajectories

GeoSPRINT is a training‑free framework that constructs non‑uniform sampling schedules for diffusion model inference by detecting geometrically redundant steps in denoising trajectories. It uses a hyperplanarity test in latent space, implemented via QR factorization, to allocate more steps to high‑curvature regions, and introduces the trajectory projection score α_traj as a model‑free diagnostic for flow quality. Across CIFAR‑10, LSUN Church, and Stable Diffusion v1.5, GeoSPRINT consistently outperforms uniform DDIM schedules at matched NFE budgets, improving FID scores by up to 1.93 points.

By Arpita Joshi