arXiv Machine Learning

Estimating the Causal Effects of T Cell Receptors

The paper introduces a method for estimating the causal effects of T cell receptor (TCR) sequences on patient outcomes using observational TCR sequencing and clinical data. It corrects for unobserved confounders by leveraging the pre-selection TCR repertoire generated through V(D)J recombination as a natural experiment, and employs permutation‑invariant neural networks to scale to millions of sequences. The approach is validated on semisynthetic data and applied to COVID‑19 severity, identifying TCRs that are observed in patients, bind SARS‑CoV‑2 antigens in vitro, and positively influence clinical outcomes.

arXiv AI
Jul 14

CDFM: Towards a General-Purpose Causal Discovery Foundation Model

arXiv:2607. 11508v1 Announce Type: cross Abstract: Causal discovery, the process of recovering underlying causal structures from observational data, is a fundamental pursuit across scientific disciplines.

By Jie Qiao, Ruichu Cai, Zijian Li, Weilin Chen, Pengfei Hua, Boyan Xu, Zhengming Chen, Zhifeng Hao, Peng Cui
arXiv Machine Learning
Jul 23

SwiftRepertoire: Few-Shot Immune-Signature Synthesis via Dynamic Kernel Codes

arXiv:2602. 01051v5 Announce Type: replace Abstract: Repertoire-level analysis of T cell receptors offers a biologically grounded signal for disease detection and immune monitoring, yet practical deployment is impeded by label sparsity, cohort heterogeneity, and the computational burden of adapting large encoders to new tasks.

By Rong Fu, Muge Qi, Yang Li, Yabin Jin, Jiekai Wu, Chunlei Meng, Juntao Gao, Li Bao, Qi Zhao, Wei Luo, Youjin Wang, Simon Fong
Hugging Face Trending Papers
Jul 22

Local Causal Structure Learning in the Presence of Latent Variables and Selection Bias

Discovering the direct causes and effects of a target variable from observational data is a fundamental problem in causal discovery, with broad applications in domains such as gene regulatory analysis and biomedical research. Existing causal discovery methods either learn a global causal structure, which incurs substantial computational cost, or assume the absence of latent variables and selection bias, assumptions that are often violated in real-world settings.

arXiv Machine Learning
2d ago

CIDER-FM: Foundation Models for Causal Inference from Diverse Experimental Regimes

CIDER-FM is a causal foundation model that combines finite observational data with surrogate-interventional datasets to predict target conditional interventional distributions more accurately than using observational data alone. It employs an intervention-aware representation and hierarchical three‑axis attention to integrate information across variables, samples, and experimental regimes. Experiments on synthetic graphs, simulated data, and real‑world Causal Chambers data show that incorporating experimental context improves CID prediction performance.

By Yuche Gao, Arik Reuter, Siyuan Guo, Anish Dhir, Bernhard Sch\"olkopf, Adrian Weller
arXiv Machine Learning
Sep 15

An immune world model for multiscale forecasting and therapeutic hypothesis generation

arXiv:2609.14709v1 Announce Type: new Abstract: Immune therapies act across cell-intrinsic programs, tissue ecosystems, and patient-specific immune states, yet most predictors address these scales se...

By Taoyong Cui, Xi Wang, Zonghang Li, Jinchao Ding, Lingsen You, Yuzhi Xu, Wanghan Xu, Fang Wu, Kejun Ying, Wanli Ouyang, Pheng Ann Heng, Ling Yang, Zhenfei Yin, Yingcheng Wu
arXiv AI
4d ago

Explainability from Training with Applications to TCR-Epitope Prediction

The paper introduces Explainability from Training (EFT), a model‑agnostic method that tracks how deep learning models learn and organize evidence during training. EFT is applied to four leading T cell receptor‑epitope prediction models, revealing distinct learning trajectories for CNNs and transformers, conflicts between TCR alpha and beta chain evidence, and differences in feature preferences when using real versus predicted structural data. The authors also present a new benchmark, TCR‑XAI2, comprising 388 experimentally resolved TCR‑epitope structures and several predicted models to evaluate these insights.

By Jiarui Li, Zixiang Yin, Samuel Landry, Zhengming Ding, Ramgopal Mettu
arXiv Machine Learning
Jun 9

Causal Representation Learning from Network Data

arXiv:2509. 01916v2 Announce Type: replace Abstract: Causal disentanglement from soft interventions is identifiable under the assumptions of linear interventional faithfulness and availability of both observational and interventional data.

By Jifan Zhang, Michelle M. Li, Elena Zheleva