arXiv Machine Learning

Copula Adapted Directed Acyclic Graph for Cluster Representation of Biomedical Data

The paper presents Copula Adapted Directed Acyclic Graph (CopDAG), a framework that combines copula models with an ensemble of causal structure discovery methods based on Directed Acyclic Graphs to represent biomedical data. By capturing non‑Gaussian, non‑linear dependencies and stable causal relationships, CopDAG enables clustering of unlabeled biomedical data using K‑means. Across 16 biomedical datasets, CopDAG achieves the highest normalized clustering accuracy and adjusted Rand index among 12 evaluated methods, and it can predict class labels and provide explainable causal visualizations without relying on data annotations.

arXiv Machine Learning
Jun 5

Binary Gaussian Copula Synthesis: an LLM-powered data augmentation framework for early dialysis prediction in chronic kidney disease

arXiv:2403. 00965v2 Announce Type: replace-cross Abstract: Only a small fraction of patients with chronic kidney disease (CKD) progress to dialysis, creating severe class imbalance that limits the performance of machine learning models for early dialysis prediction.

By Hamed Khosravi, Milad Khanchi, Mobina Noori, Srinjoy Das, Abdullah Al-Mamun, Imtiaz Ahmed
Hugging Face Trending Papers
Jul 27

Localized Anomaly Detection via Differentiable D-vine Copulas

Vine copulas provide a flexible framework for modeling complex multivariate distributions through a hierarchical decomposition into bivariate pair-copulas. Fitting a D-vine requires selecting a copula family and parameter configuration for each pair-copula from a set of candidates encoding different dependence patterns.

arXiv AI
Jul 7

Towards Diverse and Comprehensive Benchmarks for Mutual Information Estimation

arXiv:2607. 03487v1 Announce Type: cross Abstract: Mutual information (MI) estimation is a central problem in machine learning and statistics; however, existing benchmarks typically evaluate estimators on simplified, low-dimensional distributions, leaving their performance on complex, realistic data largely unexplored.

By Alberto Foresti, Ivan Butakov, Alexander Tolmachev, Giulio Franzese, Alexey Frolov, Pietro Michiardi
arXiv Machine Learning
Sep 16

Explainable Graph-theoretical Machine Learning with Application to Alzheimer's Disease Prediction

The paper introduces Explainable Graph-theoretical Machine Learning (XGML) to build individual metabolic brain graphs from FDG-PET data and identify subgraphs predictive of multivariate Alzheimer’s disease outcomes. Using ADNI data, the best model—kernel density estimation with Hellinger distance and random forest—achieved a Pearson correlation of 0.595 across eight cognitive scores, with the highest performance on ADAS13, ADAS11, and ADASQ4. Key edges were found to be jointly but differentially predictive, indicating potential network biomarkers for cognitive decline, though external validation on OASIS3 showed weaker performance likely due to cohort differences.

By Narmina Baghirova, Duy-Thanh V\~u, Duy-Cat Can, Christelle Schneuwly Diaz, Julien Bodlet, Guillaume Blanc, Georgi Hrusanov, Bernard Ries, Oliver Y. Ch\'en
arXiv AI
Aug 28

Methodological and Conceptual Framework for 5D Multi-Table Analysis: A Unified Approach for Complex Data Reuse

The paper introduces the Relational Hypergraph Transformer (RHT), a unified architecture that models relational databases as hypergraphs and learns pentadimensional embeddings (PentE). RHT applies sparse relational attention whose complexity scales with the average relational degree, making it computationally efficient for large, high‑dimensional, and high‑cardinality datasets. Experiments on the Synthea synthetic electronic health record dataset show that RHT produces more semantically coherent embeddings than tabular, relational, and temporal graph baselines, while remaining scalable, and the authors provide an open‑source implementation and plan clinical validation on MIMIC‑IV.

By Edouard Lansiaux, Hugo Kazzi, Aur\'elien Loison, Slim Hammadi, Emmanuel Chazard