arXiv:2606. 09558v1 Announce Type: cross Abstract: Motivation: Transformer-based models are increasingly applied to large-scale single-cell transcriptomics, showing strong performance through self-supervised learning on millions of cells.
By Mikele Milia, Louis Fabrice Tshimanga, Henning Mueller, Manfredo Atzori, Barbara Di Camillo
The paper introduces scTrilemma, a latent-bottleneck variational autoencoder designed to address the representation trilemma in single‑cell RNA‑seq data: preserving biological identity and state, remaining robust to nuisance context, and retaining gene‑level variation for expression analysis. scTrilemma routes expression‑derived variation to the embedding, decoder, or prior, gating gene tokens by expression and conditioning the prior on unlabeled pseudo‑bulk context, all under a single reconstruction objective without target annotations. In zero‑shot evaluations on successive CZ CELLxGENE Census releases, scTrilemma simultaneously satisfies all three demands, maintaining biological state, differential‑expression, and pathway structure across multiple disease settings, and latent interventions show context can be removed with minimal impact on other demands.
By Yunhak Oh, Yoonho Lee, Junseok Lee, Namkyeong Lee, Sang-Yeon Hwang, Yinhua Piao, Hyomin Kim, Seonghwan Kim, Jaechang Lim, Woo Youn Kim, Sungsoo Ahn, Chanyoung Park
arXiv:2608. 00985v1 Announce Type: new Abstract: The rapid growth of single-cell transcriptomic data has enabled the development of foundation models pretrained primarily by reconstructing masked expression values.
By Jiaqi Xiong, Yuntao hu, Yu Zheng, Yifei Shi, Xinyue Guo, Jiaxin Qi
arXiv:2606. 14734v1 Announce Type: cross Abstract: Motivation: Gene regulatory network inference from single-cell RNA sequencing (scRNA-seq) data is important for uncovering cell-state-specific transcriptional programs.
By Ziyang Dong, Shanwen Tan, Hengchuang Yin, Wei Liu, Yifan Wang, Siyu Yi, Jiancheng Lv, Wei Ju
arXiv:2606. 00685v1 Announce Type: new Abstract: Gene regulatory networks (GRNs) capture transcription factor-target interactions and are central to understanding cell-state regulation and disease.
By Tianyang Xu, Tianci Liu, Niraj Rayamajhi, Ryan Patrick, Kranthi Varala, Ying Li, Jing Gao
CRNDiff is a new count‑native diffusion framework that uses stochastic chemical reaction networks to model nonnegative integer data such as single‑cell RNA sequencing. It provides a closed‑form forward‑noising kernel, enabling efficient reverse sampling via forward‑filtering backward‑sampling and data‑driven selection of the terminal noising time. The method also introduces tilted Feynman–Kac steering to sample rare subpopulations without retraining, and demonstrates superior conditional fidelity and marker‑level preservation on human heart scRNA‑seq data.
By Yuxuan Qiu, Praful Gagrani, Tetsuya J Kobayashi
PopPert is a framework that models population-level joint gene expression distributions to predict transcriptional responses to perturbations in single-cell RNA sequencing data. By using a low‑rank Gaussian Copula, it captures gene co‑expression patterns and eliminates the need for cell‑to‑cell correspondence, thereby reducing sensitivity to single‑cell noise. Across multiple benchmarks, PopPert outperforms existing methods in differential expression recovery, perturbation effect estimation, and distribution matching, demonstrating the effectiveness of population‑level joint distribution learning for unpaired single‑cell data.
By Handong Wang, Jiaxin Qi, Haochen Feng, Baisheng Lai
arXiv:2607. 23821v1 Announce Type: new Abstract: Identifying therapeutic target genes from single-cell RNA sequencing (scRNA-seq) data remains a fundamental challenge in translational biology.
By Shuyu Chen, Chen Zhu, Ye Zhang, Yang Li, Qiqi Xie, Haohan Wang
arXiv:2606. 07760v1 Announce Type: new Abstract: Understanding cellular phenotypes and how they respond to perturbations is critical for disease biology and therapeutic design.
By Alma Andersson, Aya Abdelsalam Ismail, Edward De Brouwer, Doron Haviv, Tommaso Biancalani, Kyunghyun Cho, Gabriele Scalia, A\"icha BenTaieb, Hector Corrada Bravo
The paper introduces SCR-MF, a two‑stage workflow for single‑cell RNA sequencing imputation that first detects dropout events with scRecover and then imputes missing values using the non‑parametric missForest algorithm. Benchmarking on public and simulated datasets shows that SCR‑MF delivers robust, interpretable results that match or surpass existing methods while maintaining biological fidelity. Runtime analysis indicates that SCR‑MF balances accuracy with computational efficiency, making it well suited for mid‑scale single‑cell studies.
By Ali Anaissi, Deshao Liu, Yuanzhe Jia, Weidong Huang, Widad Alyassine, Junaid Akram
CellMSA introduces a novel single‑cell representation learning framework that leverages a multiple‑sequence‑alignment‑inspired context model. For each target cell, it retrieves relevant cells across batches and related cell types, summarizing cross‑cell patterns into a context‑dependent gene‑pair representation that is fed into a pair‑aware encoder. Pretraining on a massive human single‑cell corpus (≈109 million cells) and subsequent benchmarks demonstrate consistent performance gains over existing methods.
By Suyuan Zhao, Minghao Liu, Yizhen Luo, Zaiqing Nie
arXiv:2606. 07676v1 Announce Type: cross Abstract: Spatial transcriptomics (ST) is a powerful tool for exploring biological properties dependent on structure, proximity, and interaction in tissue.
By Joseph Boyd, Matthew Lyon, Martino Mansoldo, Christian Hurry, Finnian Firth