A Theoretical Framework for Statistical Evaluability of Generative Models
arXiv:2604. 05324v2 Announce Type: replace Abstract: Statistical evaluation aims to estimate the generalization performance of a model using held-out i.
arXiv:2604. 05324v2 Announce Type: replace Abstract: Statistical evaluation aims to estimate the generalization performance of a model using held-out i.
arXiv:2501. 18897v4 Announce Type: replace-cross Abstract: Generative models have achieved remarkable success across a range of applications, yet their evaluation still lacks principled uncertainty quantification.
arXiv:2606. 23872v1 Announce Type: cross Abstract: As generative models increasingly produce samples that are indistinguishable from human-created content, it becomes difficult to determine whether a given data point was part of a model's natural training set or was generated by the model itself, especially when models memorize and reproduce training data.
arXiv:2603.08064v3 Announce Type: replace Abstract: Most evaluations of generative models rely on feature-distribution metrics such as FID, which operate on continuous recognition features that are e...
arXiv:2511. 02414v3 Announce Type: replace Abstract: With the recent success of generative models in image and text, the question of their evaluation has recently gained a lot of attention.
arXiv:2607. 04360v1 Announce Type: cross Abstract: Conditional generative models have emerged as powerful tools for sampling from target conditional distributions, driving substantial advances across a wide range of scientific and applied domains.
arXiv:2608. 09117v1 Announce Type: new Abstract: Probabilistic Circuits (PCs) are tractable generative models whose internal nodes encode a hierarchy of probabilistic sum- maries over different variable scopes.
arXiv:2608.24492v1 Announce Type: cross Abstract: Uncertainty quantification (UQ) methods are widely used for hallucination detection in large language models (LLMs) in closed-book settings where gro...
Uncertainty quantification (UQ) methods are widely used for hallucination detection in large language models (LLMs) in closed-book settings where ground-truth evidence is unavailable at inference time...
arXiv:2606. 19496v1 Announce Type: new Abstract: Generative models can produce individually plausible samples while deviating substantially from a target set in the distribution of key features.
arXiv:2608.31117v1 Announce Type: cross Abstract: Generative models have become central across science and industry, from image and text synthesis to the design of molecules and materials. Quantum ge...
The paper introduces techniques for measuring the robustness of predictions made by two generative classifiers—naive Bayes classifiers and generative forests—whose underlying models are probabilistic graphical models. Robustness is defined as the degree to which the classifier’s distribution can be perturbed without altering its prediction, with perturbations explored via epsilon‑contamination, total variation distance, and chi‑squared divergence neighborhoods. Experiments on benchmark datasets show that the computed robustness values can serve as indicators of prediction trustworthiness and are compared against other existing indicators.