arXiv AI

Towards Diverse and Comprehensive Benchmarks for Mutual Information Estimation

arXiv:2607. 03487v1 Announce Type: cross Abstract: Mutual information (MI) estimation is a central problem in machine learning and statistics; however, existing benchmarks typically evaluate estimators on simplified, low-dimensional distributions, leaving their performance on complex, realistic data largely unexplored.

arXiv Statistics ML
Sep 3

Copula Transformations for Data-Consistent Inversion

The paper introduces a copula-based framework to relate Data‑Consistent Inversion (DCI) and its iterative variant (iDCI). By applying Sklar’s theorem, the authors factor the DCI update into marginal and dependence components, showing that any remaining discrepancy after iDCI convergence is fully captured by the copulas of the observed and predicted joint distributions. They prove that an exact copula transformation recovers the original DCI solution and provide convergence results for approximate transformations, supported by numerical examples illustrating adaptive refinement and progressive problem refinement.

By Troy Butler, Tianyi Jiang, Jo\~ao Silva, Harri Hakula, Timothy Wildey
arXiv Machine Learning
4d ago

Gaussian Mixture Copula Processes for Irregular Time Series

arXiv:2605.23632v2 Announce Type: replace Abstract: We introduce Gaussian Mixture Copula Processes (GMCP), a conditional copula process for irregularly sampled multivariate time series (IMTS) that is...

By Christian Kl\"otergens, Tom Hanika, Lars Schmidt-Thieme, Vijaya Krishna Yalavarthi
arXiv Machine Learning
Sep 16

Copula Adapted Directed Acyclic Graph for Cluster Representation of Biomedical Data

The paper presents Copula Adapted Directed Acyclic Graph (CopDAG), a framework that combines copula models with an ensemble of causal structure discovery methods based on Directed Acyclic Graphs to represent biomedical data. By capturing non‑Gaussian, non‑linear dependencies and stable causal relationships, CopDAG enables clustering of unlabeled biomedical data using K‑means. Across 16 biomedical datasets, CopDAG achieves the highest normalized clustering accuracy and adjusted Rand index among 12 evaluated methods, and it can predict class labels and provide explainable causal visualizations without relying on data annotations.

By Heranga K. Rathnasekara, Norou Diawara, Manar D. Samad
Hugging Face Trending Papers
Jul 27

Localized Anomaly Detection via Differentiable D-vine Copulas

Vine copulas provide a flexible framework for modeling complex multivariate distributions through a hierarchical decomposition into bivariate pair-copulas. Fitting a D-vine requires selecting a copula family and parameter configuration for each pair-copula from a set of candidates encoding different dependence patterns.

arXiv Machine Learning
4d ago

ALICE: In-context, Zero-shot, Mutual Information Estimation

ALICE is a foundation model that estimates mutual information (MI) without per‑distribution training. Trained only on synthetic distributions, it acts as an in‑context estimator of rectified‑flow velocity fields, producing MI via a fixed identity that integrates squared differences between joint and conditional fields. The authors validate ALICE on a challenging benchmark and demonstrate its applicability to unseen data in biology, genetics, and neuroscience, achieving performance comparable to neural estimators trained separately for each distribution.

By Giulio Franzese, Simone Rossi, Pietro Michiardi
arXiv AI
Aug 20

From Inference to Adaptation: A Unified Optimal Transport View of Vision Language Model

arXiv:2608. 18339v1 Announce Type: cross Abstract: Vision-language models (VLMs) have demonstrated remarkable zero-shot capabilities yet remain sensitive to real-world distribution shifts during inference.

By Qi Yu, Zhichen Zeng, Katherine Tieu, Xiyuan Yang, Ruizhong Qiu, Yuchen Yan, Lihui Liu, Yanjun Zhao, Lingjie Chen, Jingrui He, Hanghang Tong
arXiv Machine Learning
Jun 5

Inverse Entropic Optimal Transport Solves Semi-supervised Learning via Data Likelihood Maximization

arXiv:2410. 02628v5 Announce Type: replace Abstract: Learning conditional distributions $\pi^*(\cdot|x)$ is a central problem in machine learning, which is typically approached via supervised methods with paired data $(x,y) \sim \pi^*$.

By Mikhail Persiianov, Arip Asadulaev, Nikita Andreev, Nikita Starodubcev, Dmitry Baranchuk, Anastasis Kratsios, Evgeny Burnaev, Alexander Korotin