arXiv Machine Learning

Scalable partial information decomposition for symptom networks via supervised embeddings

The paper introduces ePID, an embedding-based approach that scales partial information decomposition (PID) to large symptom networks by compressing non‑focal symptoms into a low‑cardinality discrete embedding. Using a supervised Agglomerative Conditional Information Bottleneck (ACIB) embedding, ePID accurately recovers source‑unique, remainder‑unique, redundant, and synergistic components for each ordered source‑target pair across 83 real‑world datasets, outperforming 12 other embeddings. Applied to PHQ‑9 and the Interpersonal Reactivity Index, ePID reveals distinct patterns of redundancy and synergy that align with each instrument’s construction, demonstrating its ability to separate overlapping from interaction‑dependent information in symptom networks.

arXiv Machine Learning
Sep 3

Omega-N: Interpretable Structural Node Descriptors and Their Applicability Domain

The paper introduces Omega‑N, a set of ten interpretable node‑level structural descriptors derived from localizing four factors of a composite structural index. By correcting the ill‑conditioned localization with a configuration‑null excess and a multi‑scale personalized‑PageRank neighbourhood, Omega‑N achieves competitive or superior performance in six in‑domain node‑classification tasks compared to a recursive feature engine that uses up to 252 features. In drug‑target prioritisation on protein interaction networks, Omega‑N improves AUPRC by 0.073 to 0.144 over a centrality baseline and remains robust across independent datasets and bias controls, though it offers no benefit when combined with Node2Vec. whyItMatters":"The study demonstrates that a compact, interpretable set of structural features can match or exceed more complex feature sets in practical graph‑based prediction tasks, particularly in biomedical network analysis."

By Alberto Acedo
arXiv Machine Learning
Jul 21

Differentiable latent structure discovery for interpretable forecasting in clinical time series

arXiv:2604. 27967v2 Announce Type: replace Abstract: Background: We introduce StructGP, a continuous-time multi-task Gaussian process that couples process convolutions with differentiable structure learning to uncover a sparse, ordered directed acyclic graph (DAG) of inter-variable dependencies while preserving principled uncertainty.

By Ivan Lerner, Jean Feydy, Alexandre Kalimouttou, Anita Burgun, Francis Bach
arXiv AI
Jun 12

Is It You or Your Environment? A Bayesian Inference Framework for Genomically-Anchored Personalized Physiological Interpretation

arXiv:2606. 13556v1 Announce Type: new Abstract: Personalized health AI systems face a fundamental cold-start problem: machine learning models for physiological interpretation require weeks of individual behavioral data before they can distinguish constitutional variation from environmentally driven deviation.

By Aruna Dey, Suraj Biswas
arXiv Machine Learning
Jun 18

Shrinkage priors for Bayesian Substitute Confounders

arXiv:2606. 18535v1 Announce Type: cross Abstract: Multi-cause observational studies contain information about unmeasured confounding through the dependence structure among causes.

By Yordan P. Raykov, Hengrui Luo, Justin D. Strait, Wasiur R. KhudaBukhsh
arXiv AI
Jun 10

Dep-LLM: Training-Free Depression Diagnosis via Evidence-Guided Structured Multi-factor with Reliable LLM Reasoning

arXiv:2606. 10796v1 Announce Type: cross Abstract: Automatic Depression Detection (ADD) from clinical interviews is a pivotal task in computational mental health, yet it remains challenging due to two critical obstacles: 1) difficulty in modeling complex but sparsely distributed depression clues within lengthy, multi-topic clinical interviews, leading to superficial and unreliable reasoning; 2) scarcity of labeled data due to clinical privacy, together with high cost of training and fine-tuning, limiting the deployment of supervised ADD systems.

By Yiqing Lyu, Xianbing Zhao, Buzhou Tang, Ronghuan Jiang
arXiv AI
Sep 17

Information Set Emulation: Causal Certificates for AI Derived EHR Features

The paper introduces information set emulation, a method that attaches detailed causal certificates—such as source evidence, timing, and proposed causal roles—to AI‑derived features extracted from electronic health records (EHRs). These certificates provide auditable evidence for causal roles and guide whether a feature can be used for causal inference or should be routed to compatible reporting or separate analyses. The framework integrates with a joint EHR observation map and offers identification, estimation, and diagnostic tools under standard causal assumptions, illustrated through synthetic simulations and a finite‑world example.

By Takes Fujita (VRI), Nobutaka Hattori (Department of Neurology, Juntendo University School of Medicine)