arXiv AI By Suyuan Zhao, Minghao Liu, Yizhen Luo, Zaiqing Nie

CellMSA: Context Modeling for Single-Cell Representation Learning

Read the original on arXiv AI →

CellMSA introduces a novel single‑cell representation learning framework that leverages a multiple‑sequence‑alignment‑inspired context model. For each target cell, it retrieves relevant cells across batches and related cell types, summarizing cross‑cell patterns into a context‑dependent gene‑pair representation that is fed into a pair‑aware encoder. Pretraining on a massive human single‑cell corpus (≈109 million cells) and subsequent benchmarks demonstrate consistent performance gains over existing methods.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 3

Subcellularly Resolved Single-Cell Embedding Learning with Transcriptomic data, Protein Structure and Localization Information

The paper introduces a multimodal framework that learns subcellularly resolved cell embeddings by integrating RNA expression profiles, protein sequence representations, and protein structural information using a cross‑attention architecture. This approach models interactions within distinct subcellular compartments, producing fine‑grained embeddings that capture both molecular expression patterns and functional protein properties. It is presented as the first method to jointly incorporate transcriptomic data, sequence, and structural knowledge for subcellularly resolved cell representation.

By Zhen Zhou, Jiachen Li, Yuan Liu, Xiaoyong Pan, Hong-Bin Shen
arXiv AI
2d ago

scTrilemma: Balancing Identity, Invariance, and Fidelity in Single-Cell Representation Learning

The paper introduces scTrilemma, a latent-bottleneck variational autoencoder designed to address the representation trilemma in single‑cell RNA‑seq data: preserving biological identity and state, remaining robust to nuisance context, and retaining gene‑level variation for expression analysis. scTrilemma routes expression‑derived variation to the embedding, decoder, or prior, gating gene tokens by expression and conditioning the prior on unlabeled pseudo‑bulk context, all under a single reconstruction objective without target annotations. In zero‑shot evaluations on successive CZ CELLxGENE Census releases, scTrilemma simultaneously satisfies all three demands, maintaining biological state, differential‑expression, and pathway structure across multiple disease settings, and latent interventions show context can be removed with minimal impact on other demands.

By Yunhak Oh, Yoonho Lee, Junseok Lee, Namkyeong Lee, Sang-Yeon Hwang, Yinhua Piao, Hyomin Kim, Seonghwan Kim, Jaechang Lim, Woo Youn Kim, Sungsoo Ahn, Chanyoung Park