What FID Hides: Detecting, Ranking, and Diagnosing Deviations in Generative Evaluation
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2608. 07713v1 Announce Type: cross Abstract: Latent medical image generators usually treat the tokenizer as fixed preprocessing.
DANCo (Dimensionality from Angle and Norm Concentration) jointly calibrates nearest-neighbor distance and angular statistics and consistently reaches state-of-the-art accuracy on clean intrinsic-dimen...
arXiv:2609.37114v1 Announce Type: new Abstract: DANCo (Dimensionality from Angle and Norm Concentration) jointly calibrates nearest-neighbor distance and angular statistics and consistently reaches s...
arXiv:2409. 10094v3 Announce Type: replace-cross Abstract: Out-of-Distribution (OoD) detection aims to justify whether a given sample is from the training distribution of the classifier-under-protection, i.
The paper investigates the use of the squared norm of a whitened foundation‑model embedding as a training‑free likelihood surrogate. It shows that the apparent Gaussianity of whitened coordinates stems from the projection central limit theorem, not from a true joint Gaussian distribution, and that the norm is systematically over‑dispersed compared to a Gaussian reference. The authors explain that whitening reverses the encoder’s spectral hierarchy, concentrating norm contributions in near‑degenerate directions dominated by noise, and propose interpreting the squared norm as a Mahalanobis measure of semantic atypicality rather than a log‑likelihood.
The paper introduces D3-Omni, a balanced and decoupled benchmark designed to diagnose fine‑grained multimodal understanding in OmniJudges that evaluate text‑to‑image, text‑to‑video, and text‑to‑speech generation. D3-Omni covers 53 orthogonal binary dimensions across 10,671 samples, using fixed positive seeds and controlled prompt rewriting to generate negatives, thereby ensuring each error can be attributed to a single capability. The benchmark’s dual‑balanced, decoupled, and dynamic design achieves near 1:1 per‑dimension parity and a uniform total‑score distribution, revealing that strong OmniJudges often miss modality‑related failures and treat distinct attributes as a single decision, masking systematic blind spots.