arXiv Machine Learning

Linear and Quadratic Discriminant Analysis: Tutorial

Towards Data Science
Sep 6

Linear Discriminant Analysis (LDA) in Real-Life: Dimensionality Reduction in a Real-Estate Dataset

The article discusses applying Linear Discriminant Analysis (LDA) to reduce dimensionality in a real‑estate dataset for classification tasks. It explains how LDA can transform high‑dimensional data into a lower‑dimensional space while preserving class separability. The post demonstrates the practical use of LDA in a real‑life scenario, specifically within the real‑estate domain.

By Carolina Bento
arXiv Machine Learning
Sep 22

On Generalized Naive Bayes with Continuous Features

The paper extends the Generalized Naive Bayes (GNB) model to handle continuous explanatory variables. It shows that GNB structure learning depends only on pair copulas of bivariate marginals and can be framed as a matroid, enabling greedy algorithms that minimize Kullback–Leibler divergence. Three model variants are explored—joint Gaussian, Gaussian copula with arbitrary marginals, and fully arbitrary copula and marginals—along with a GNB forest-based model reduction method and empirical comparisons to classical glass‑box classifiers.

By \'Abrah\'am Papp, Botond Szil\'agyi, Edith Alice Kov\'acs
arXiv Machine Learning
Aug 21

Exact Algebraic Computation of Learning Coefficients for Two-Dimensional Singular Models

arXiv:2608. 20183v1 Announce Type: new Abstract: Classical information criteria such as the Bayesian Information Criterion (BIC) rely on regularity assumptions that break down for singular models, leading to incorrect model selection in settings such as deep learning.

By Gr\'egoire Sergeant-Perthuis (CQSB, Sorbonne Universit\'e), Elias Tsigaridas (Ouragan Team, INRIA), Jules Tsukahara (Ouragan Team, INRIA)
arXiv Machine Learning
Sep 23

Learning to bin: differentiable and Bayesian optimization for multi-dimensional discriminants in high-energy physics

The paper introduces a method for optimizing bin boundaries in multi-dimensional discriminants using a Gaussian Mixture Model (GMM), allowing flexible definition of analysis categories. Two optimization strategies—differentiable and Bayesian—are compared in toy binary and three-class setups, with the differentiable approach excelling in multi-dimensional cases. Applied to the FAIR Universe $H ightarrow au au$ dataset, the GMM-based optimization achieves the highest signal significance, and the tools are released as lightweight Python plugins.

By Johannes Erdmann, Nitish Kumar Kasaraguppe, Florian Mausolf
arXiv Machine Learning
Sep 10

Learning Kernels by Alignment for Multiclass Bayes Classification

The paper introduces a framework that learns the kernel used in kernel methods through alignment, leveraging the Collaborative Learning and Inference (CLaI) approach. It demonstrates that CLaI can be interpreted as a kernel alignment process and that its inference stage is equivalent to kernel Bayes classification with Parzen-window density estimation. By replacing cosine similarity with a learned Mahalanobis distance, the authors extend CLaI to multiclass classification, achieving higher accuracy, faster convergence, and lower calibration error on datasets such as CIFAR-10, PathMNIST, and SleepEDF, while also showing connections to Gaussian processes and competitive calibration in sepsis prediction.

By Hollan Haule, Alfredo Gonzalez-Sulser, Javier Escudero