arXiv AI

scLLM-DSC: LLM-Knowledge Enhanced Cross-Modal Deep Structural Clustering for Single-Cell RNA Sequencing

arXiv:2606. 13007v1 Announce Type: cross Abstract: Clustering is fundamental to scRNA-seq analysis, serving as a cornerstone for identifying cell populations and resolving tissue heterogeneity.

arXiv AI
Jun 18

scGTN: Deep Siamese Graph Transformer Network for Single-cell RNA Sequencing Clustering

arXiv:2606. 18672v1 Announce Type: cross Abstract: Single-cell RNA sequencing (scRNA-seq) serves a pivotal role in characterizing gene expression at the cellular level, enabling the identification of cell types and advancing the understanding of cellular heterogeneity.

By Jinke Wu, Yifan Wang, Siyu Yi, Caiyang Yu, Ziyue Qiao, Nan Yin, Jiancheng Lv, Wei Ju
arXiv AI
3d ago

CellMSA: Context Modeling for Single-Cell Representation Learning

CellMSA introduces a novel single‑cell representation learning framework that leverages a multiple‑sequence‑alignment‑inspired context model. For each target cell, it retrieves relevant cells across batches and related cell types, summarizing cross‑cell patterns into a context‑dependent gene‑pair representation that is fed into a pair‑aware encoder. Pretraining on a massive human single‑cell corpus (≈109 million cells) and subsequent benchmarks demonstrate consistent performance gains over existing methods.

By Suyuan Zhao, Minghao Liu, Yizhen Luo, Zaiqing Nie
arXiv AI
Sep 3

Subcellularly Resolved Single-Cell Embedding Learning with Transcriptomic data, Protein Structure and Localization Information

The paper introduces a multimodal framework that learns subcellularly resolved cell embeddings by integrating RNA expression profiles, protein sequence representations, and protein structural information using a cross‑attention architecture. This approach models interactions within distinct subcellular compartments, producing fine‑grained embeddings that capture both molecular expression patterns and functional protein properties. It is presented as the first method to jointly incorporate transcriptomic data, sequence, and structural knowledge for subcellularly resolved cell representation.

By Zhen Zhou, Jiachen Li, Yuan Liu, Xiaoyong Pan, Hong-Bin Shen
Hugging Face Trending Papers
Sep 2

Subcellularly Resolved Single-Cell Embedding Learning with Transcriptomic data, Protein Structure and Localization Information

The paper introduces a multimodal framework that learns subcellularly resolved cell embeddings by integrating RNA expression profiles, protein sequence representations, and protein structural information. It uses a cross‑attention architecture to model interactions across distinct subcellular compartments, producing embeddings that capture both molecular expression patterns and functional protein properties. This approach is presented as the first to jointly incorporate transcriptomic data, protein sequences, and structural knowledge within a unified cross‑modal learning paradigm.

arXiv AI
Jun 30

Data-Efficient Multimodal Alignment for Histopathology-based Molecular Prediction

arXiv:2606. 29949v1 Announce Type: cross Abstract: H&E-stained whole-slide images offer cohort-scale availability and rich spatial context but lack molecular specificity, whereas bulk RNA-seq provides transcriptome-wide resolution at high cost with limited archival availability.

By Dominik Winter, Dominik Vonficht, Lo\"ic Le Bescond, Christian Gebbe, Marco Rosati, Richard J. Chen, Markus Schick, Ross Stewart, Nicolas Brieu
arXiv AI
6d ago

LapDDPM: Spectral Perturbation Diffusion for Robust Single-Cell Manifold Generation

LapDDPM is a conditional Graph Diffusion Probabilistic Model that generates high‑fidelity, biologically plausible single‑cell RNA sequencing data. It incorporates graph‑based inductive biases and a spectral adversarial perturbation mechanism to enforce robustness against structural noise, effectively acting as a Distributionally Robust Optimization framework. The model extends to spatial transcriptomics and multi‑modal data, and experimental results on datasets such as PBMC3K, Dentate Gyrus, HLCA, Visium, and 10x Multiome show it outperforms state‑of‑the‑art baselines in distribution matching, manifold preservation, and downstream utility.

By Lorenzo Bini, Stephane Marchand-Maillet
arXiv Machine Learning
Sep 25

Selective Inference for Deep Clustering in Latent Spaces

The paper introduces a selective inference framework tailored for deep clustering that uses a fixed pretrained encoder to map high‑dimensional data into a latent space before clustering. It addresses the complex selection bias arising from the nonlinear transformation and offers a computationally tractable method to perform valid statistical tests on cluster differences. Experiments on synthetic data show controlled Type I error and higher power compared to conservative baselines, while genomic case studies demonstrate the ability to uncover significant cluster differences while properly accounting for selection bias.

By Eina Mizui, Tomohiro Shiraishi, Shunichi Nishino, Ichiro Takeuchi
arXiv AI
Aug 12

JEPA-DNA: Grounding Genomic Foundation Models through Joint-Embedding Predictive Architectures

arXiv:2602. 17162v3 Announce Type: replace Abstract: Genomic Foundation Models (GFMs) typically rely on Masked Language Modeling (MLM) or Next-Token Prediction (NTP) to learn the "Laws of Nature".

By Ariel Larey, Elay Dahan, Amit Bleiweiss, Raizy Kellerman, Guy Leib, Omri Nayshool, Dan Ofer, Tal Zinger, Dan Dominissini, Gideon Rechavi, Nicole Bussola, Simon Lee, Shane O'Connell, Dung Hoang, Marissa Wirth, Alexander W. Charney, Nati Daniel, Yoli Shavit