arXiv Machine Learning

On the Recoverability of Causal Relations from Bulk Gene Expression Data

arXiv:2606. 00568v1 Announce Type: new Abstract: Bulk gene expression profiling, which aggregates pooled RNA across cells within a biological sample, remains important in the single-cell era because it is typically less noisy, more sensitive, and more cost-effective than single-cell assays.

arXiv Statistics ML
Sep 16

Causal Path Analysis from Perturbational and Population-Scale Single-Cell Data with Multiscale Confounding and Measurement Error

The paper presents a framework that merges single‑cell perturbation experiments with population‑scale single‑cell data to perform causal path analysis of gene regulation. It incorporates externally learned ancestral relationships to constrain network topology, re‑estimates direct edges from population data, and applies a surrogate‑variable procedure plus errors‑in‑variables correction to handle multiscale heterogeneity and measurement error. The authors provide theoretical guarantees for confounder recovery and high‑dimensional estimation, and demonstrate the method’s effectiveness through simulations and an acute myeloid leukemia case study that uncovers distinct regulatory pathways linking transcriptional regulators to blast count.

By Kwangmoon Park, Hongzhe Li
arXiv AI
Sep 2

PopPert: Population-level Joint-Distribution Modeling for Single-Cell Perturbation Prediction

PopPert is a framework that models population-level joint gene expression distributions to predict transcriptional responses to perturbations in single-cell RNA sequencing data. By using a low‑rank Gaussian Copula, it captures gene co‑expression patterns and eliminates the need for cell‑to‑cell correspondence, thereby reducing sensitivity to single‑cell noise. Across multiple benchmarks, PopPert outperforms existing methods in differential expression recovery, perturbation effect estimation, and distribution matching, demonstrating the effectiveness of population‑level joint distribution learning for unpaired single‑cell data.

By Handong Wang, Jiaxin Qi, Haochen Feng, Baisheng Lai
arXiv AI
3d ago

scTrilemma: Balancing Identity, Invariance, and Fidelity in Single-Cell Representation Learning

The paper introduces scTrilemma, a latent-bottleneck variational autoencoder designed to address the representation trilemma in single‑cell RNA‑seq data: preserving biological identity and state, remaining robust to nuisance context, and retaining gene‑level variation for expression analysis. scTrilemma routes expression‑derived variation to the embedding, decoder, or prior, gating gene tokens by expression and conditioning the prior on unlabeled pseudo‑bulk context, all under a single reconstruction objective without target annotations. In zero‑shot evaluations on successive CZ CELLxGENE Census releases, scTrilemma simultaneously satisfies all three demands, maintaining biological state, differential‑expression, and pathway structure across multiple disease settings, and latent interventions show context can be removed with minimal impact on other demands.

By Yunhak Oh, Yoonho Lee, Junseok Lee, Namkyeong Lee, Sang-Yeon Hwang, Yinhua Piao, Hyomin Kim, Seonghwan Kim, Jaechang Lim, Woo Youn Kim, Sungsoo Ahn, Chanyoung Park
arXiv AI
Jul 14

CDFM: Towards a General-Purpose Causal Discovery Foundation Model

arXiv:2607. 11508v1 Announce Type: cross Abstract: Causal discovery, the process of recovering underlying causal structures from observational data, is a fundamental pursuit across scientific disciplines.

By Jie Qiao, Ruichu Cai, Zijian Li, Weilin Chen, Pengfei Hua, Boyan Xu, Zhengming Chen, Zhifeng Hao, Peng Cui