arXiv:2607. 11508v1 Announce Type: cross Abstract: Causal discovery, the process of recovering underlying causal structures from observational data, is a fundamental pursuit across scientific disciplines.
By Jie Qiao, Ruichu Cai, Zijian Li, Weilin Chen, Pengfei Hua, Boyan Xu, Zhengming Chen, Zhifeng Hao, Peng Cui
The paper introduces ePID, an embedding-based approach that scales partial information decomposition (PID) to large symptom networks by compressing non‑focal symptoms into a low‑cardinality discrete embedding. Using a supervised Agglomerative Conditional Information Bottleneck (ACIB) embedding, ePID accurately recovers source‑unique, remainder‑unique, redundant, and synergistic components for each ordered source‑target pair across 83 real‑world datasets, outperforming 12 other embeddings. Applied to PHQ‑9 and the Interpersonal Reactivity Index, ePID reveals distinct patterns of redundancy and synergy that align with each instrument’s construction, demonstrating its ability to separate overlapping from interaction‑dependent information in symptom networks.
By Cillian Hourican, Eric Dignum, Rick Quax, Debraj Roy
arXiv:2606. 00483v1 Announce Type: cross Abstract: Genotype-based cis-expression prediction depends on accurately modeling local regulatory architecture.
By Lei Huang, Hui Shen, Kuan-Jui Su, Chuan Qiu, Martha Isabel Gonzalez-Ramirez, Anqi Liu, Zhe Luo, Yun Gong, Yipu Zhang, Dawei Li, Chaoyang Zhang, Hong-Wen Deng
arXiv:2606. 18535v1 Announce Type: cross Abstract: Multi-cause observational studies contain information about unmeasured confounding through the dependence structure among causes.
By Yordan P. Raykov, Hengrui Luo, Justin D. Strait, Wasiur R. KhudaBukhsh
The paper introduces a method for estimating the causal effects of T cell receptor (TCR) sequences on patient outcomes using observational TCR sequencing and clinical data. It corrects for unobserved confounders by leveraging the pre-selection TCR repertoire generated through V(D)J recombination as a natural experiment, and employs permutation‑invariant neural networks to scale to millions of sequences. The approach is validated on semisynthetic data and applied to COVID‑19 severity, identifying TCRs that are observed in patients, bind SARS‑CoV‑2 antigens in vitro, and positively influence clinical outcomes.
By Eli N. Weinstein, Elizabeth B. Wood, David M. Blei
arXiv:2601. 14590v3 Announce Type: replace Abstract: Counterfactual explanations (CFEs) provide human-centric interpretability by identifying the minimal, actionable changes required to alter a machine learning model's prediction.
By Shovito Barua Soumma, Asiful Arefeen, Stephanie M. Carpenter, Melanie Hingle, Hassan Ghasemzadeh
arXiv:2609.24422v1 Announce Type: new
Abstract: Hierarchical data is ubiquitous in the empirical sciences and is most commonly analyzed with generalized linear mixed-effects models (GLMMs). Bayesian...
By Alex Kipnis, Marcel Binz, Eric Schulz
arXiv:2607. 19618v1 Announce Type: cross Abstract: Genomic language models achieve strong performance across regulatory-genomics tasks, yet what these models internally represent remains opaque, and the field lacks a principled procedure for verifying that an apparent ``concept'' inside a model is real rather than an artifact of sequence composition.
By Sarwan Ali
arXiv:2404. 02141v5 Announce Type: replace-cross Abstract: In both observational data and randomized control trials, researchers select statistical models to articulate how the outcome of interest varies with combinations of observable covariates.
By Aparajithan Venkateswaran, Anirudh Sankar, Arun G. Chandrasekhar, Tyler H. McCormick
arXiv:2606. 05797v1 Announce Type: new Abstract: Longitudinal treatment decisions require predicting potential outcomes under future treatment sequences in the presence of time-varying confounding, heterogeneous patient dynamics, and limited domain-specific data.
By Amirhossein Zare, Amirhessam Zare, Herlock Rahimi, Reza Salarikia, Mohammad Kashkooli
Hierarchical data is ubiquitous in the empirical sciences and is most commonly analyzed with generalized linear mixed-effects models (GLMMs). Bayesian inference for GLMMs yields calibrated uncertainty...
arXiv:2606. 17491v1 Announce Type: cross Abstract: Binary data factorization is common, but real-valued methods ignore discreteness and yield hard-to-interpret factors.
By Adolphus Wagala, Mehmet Samur, Giovanni Parmigiani