arXiv AI By Mehrdad Shoeibi, Niloofar Yousefi

Can We Trust In-Distribution Success? Locked Evaluation Reveals Transfer Failure and Sampling-Depth Entanglement in CRISPRi Perturbation Prediction

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv Machine Learning
Aug 31

Locked Evaluation Surfaces: Transfer Failure and Sampling-Depth Entanglement in CRISPRi Perturbation-Effect Prediction

The study evaluates a frozen Geneformer representation for predicting CRISPRi perturbation effects under a tightly controlled, pre‑registered protocol. While the representation shows significant predictive power within the Virtual Cell Challenge dataset, it fails to transfer to external screens, with negative zero‑shot Spearman correlations. The analysis also reveals that the VCC endpoint is heavily influenced by sampling depth, as cell count alone explains most of the variance, indicating a sampling‑depth entanglement that could mask transfer failures in less controlled settings.

By Mehrdad Shoeibi, Niloofar Yousefi
arXiv Machine Learning
Sep 24

Discover, Falsify, Revise: Auditing Input-Use Claims from Source Code to Predictive Contribution in Agent-Discovered Cell Models

The paper introduces CELLAUDIT, a method for auditing whether inputs claimed to influence predictive models actually do so. By testing if an input can enter the computation, whether predictions depend on it, and if that dependence improves observed responses, the authors evaluate agent-generated predictors on a morphology‑transcriptomics benchmark (BBBC047). Their findings show that many models claim compound contributions that are not supported by the data, and that falsification‑guided revisions can recover genuine input effects while improving performance.

By Mengran Li, Bo Li, Chengyang Zhang, Yang Yan, Jinfeng Xu, Zhenchao Tang
arXiv AI
Sep 2

SCALE:Scalable Conditional Atlas-Level Endpoint transport for virtual cell perturbation prediction

SCALE is a conditional transport model that treats cells as unordered sets to predict treated cell populations without requiring cell-level matching. It uses a shared set-aware encoder and a conditional DiT backbone to learn latent transport, enabling endpoint supervision that is directly delta-aligned. Across diverse perturbation types—including genetic, chemical, developmental, and immune—SCALE accurately recovers gene‑expression changes, response directions, and population structure, outperforming competing methods on CRISPR data and successfully prioritizing cytokines that elicit distinct immune responses.

By Shuizhou Chen, Lang Yu, Xueqin Lin, Xinjie Mao, Songming Zhang, Xinyu Gu, Hao Wu, Sheng Xu, Kedu Jin, Lei Bai, Quan Qian, Qin Chen, Qiang Gao, Siqi Sun, Zhangyang Gao