arXiv AI

Subcellularly Resolved Single-Cell Embedding Learning with Transcriptomic data, Protein Structure and Localization Information

The paper introduces a multimodal framework that learns subcellularly resolved cell embeddings by integrating RNA expression profiles, protein sequence representations, and protein structural information using a cross‑attention architecture. This approach models interactions within distinct subcellular compartments, producing fine‑grained embeddings that capture both molecular expression patterns and functional protein properties. It is presented as the first method to jointly incorporate transcriptomic data, sequence, and structural knowledge for subcellularly resolved cell representation.

Hugging Face Trending Papers
Sep 2

Subcellularly Resolved Single-Cell Embedding Learning with Transcriptomic data, Protein Structure and Localization Information

The paper introduces a multimodal framework that learns subcellularly resolved cell embeddings by integrating RNA expression profiles, protein sequence representations, and protein structural information. It uses a cross‑attention architecture to model interactions across distinct subcellular compartments, producing embeddings that capture both molecular expression patterns and functional protein properties. This approach is presented as the first to jointly incorporate transcriptomic data, protein sequences, and structural knowledge within a unified cross‑modal learning paradigm.

arXiv AI
2d ago

CellMSA: Context Modeling for Single-Cell Representation Learning

CellMSA introduces a novel single‑cell representation learning framework that leverages a multiple‑sequence‑alignment‑inspired context model. For each target cell, it retrieves relevant cells across batches and related cell types, summarizing cross‑cell patterns into a context‑dependent gene‑pair representation that is fed into a pair‑aware encoder. Pretraining on a massive human single‑cell corpus (≈109 million cells) and subsequent benchmarks demonstrate consistent performance gains over existing methods.

By Suyuan Zhao, Minghao Liu, Yizhen Luo, Zaiqing Nie
arXiv Machine Learning
Jun 9

Integrating gene regulatory priors into Transformer attention with scTransformer for interpretable scRNA-seq analysis

arXiv:2606. 09558v1 Announce Type: cross Abstract: Motivation: Transformer-based models are increasingly applied to large-scale single-cell transcriptomics, showing strong performance through self-supervised learning on millions of cells.

By Mikele Milia, Louis Fabrice Tshimanga, Henning Mueller, Manfredo Atzori, Barbara Di Camillo
Hugging Face Trending Papers
Aug 6

BioM-JEPA: joint-embedding prediction of graph-connected gene blocks in single cells

Single-cell transcriptomes are sparse observations of coordinated biological programmes, yet most self-supervised models learn by reconstructing individual genes. Here we present BioM-JEPA, a joint-embedding predictive architecture that instead predicts aggregate representations of graph-connected gene blocks defined by protein-association and corpus-derived coexpression evidence.

arXiv Machine Learning
Jun 5

HEIST: A Graph Foundation Model for Spatial Transcriptomics and Proteomics Data

arXiv:2506. 11152v4 Announce Type: replace-cross Abstract: Single-cell transcriptomics and proteomics have become a great source for data-driven insights into biology, enabling the use of advanced deep learning methods to understand cellular heterogeneity and gene expression at the single-cell level.

By Hiren Madhu, Jo\~ao Felipe Rocha, Tinglin Huang, Siddharth Viswanath, Smita Krishnaswamy, Rex Ying