arXiv Machine Learning By Vasily Ilin, Peter Sushko, Ranjay Krishna

DiScoFormer: Plug-In Density and Score Estimation with Transformers

Read the original on arXiv Machine Learning →

arXiv:2511. 05924v4 Announce Type: replace Abstract: Estimating probability density and its score from samples remains a core problem in generative modeling, Bayesian inference, and kinetic theory.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
3d ago

BayesNDE: Bayesian Generative Modeling for Neural Density Estimation

BayesNDE is a neural density estimator that uses Bayesian generative modeling to estimate densities without relying on invertible networks or Jacobian-determinant calculations. It constructs an adaptive proposal for each observation by inferring a sample-specific latent posterior, and then applies bridge sampling to combine proposal samples with separate posterior samples for density estimation. Experiments on synthetic datasets show improved density estimation and structure recovery, while real-world applications demonstrate better anomaly detection.

By Chenglin Li, Qiao Liu
arXiv Machine Learning
Aug 26

Generalization, memorization, and overfitting for diffusion models trained in the lazy high-dimensional regime

The paper investigates diffusion models trained in a lazy high‑dimensional regime, extending benign overfitting theory to generative settings. By analyzing denoising score matching in a vector‑valued RKHS with an inner‑product kernel, the authors derive exact risk trajectories under gradient flow when the number of samples scales proportionally with dimensionality. These trajectories reveal three distinct phases—spectral generalization, noise‑dominated interpolation, and empirical Bayes memorization—whose interplay shapes the distribution of generated samples.

By Hugo Latourelle-Vigeant, Sinho Chewi, Aram-Alexandre Pooladian, John Sous, Theodor Misiakiewicz
arXiv Machine Learning
Aug 27

Cubit: Token Mixer with Kernel Ridge Regression

The paper introduces Cubit, a Transformer‑style architecture that replaces the standard attention mechanism with Kernel Ridge Regression (KRR). By interpreting attention as Nadaraya‑Watson regression, Cubit incorporates the closed‑form KRR solution, combining kernel‑based value aggregation with normalization via the inverse kernel matrix. The authors also propose a Limited‑Range Rescale (LRR) to stabilize training and report that Cubit shows improved long‑sequence modeling, with gains increasing as training sequence length grows.

By Chuanyang Zheng, Jiankai Sun, Yihang Gao, Yuehao Wang, Liangchen Tan, Mac Schwager, Anderson Schneider, Yuriy Nevmyvaka, Xiaodong Liu
arXiv Machine Learning
1d ago

Attention Kernels for Learning Maps Between Heavy-Tailed Measures

The paper introduces attention kernels that replace the exponential function in transformer softmax to better handle operator learning on probability measures with heavy-tailed (polynomial) distributions. Two new benchmarks with closed‑form targets are constructed to evaluate how different kernel growth rates and data preprocessing affect performance. The study finds that slower‑growing kernels prevent ensemble collapse on heavy‑tailed tasks, while softmax with symlog preprocessing only succeeds on a subset of problems, and that all kernels perform similarly on Gaussian data.

By Kailen Hargenrader, Edoardo Calvello, Bohan Chen