arXiv Machine Learning

Data-Driven Soft Labeling Scales DNA Read Classification to Whole-Body Cell-Type Deconvolution

arXiv:2607. 04987v1 Announce Type: new Abstract: Cell-type deconvolution, the task of estimating the proportions of constituent cell types in a heterogeneous biological sample, is a core problem in computational biology.

arXiv Machine Learning
Jul 3

A Novel Machine Learning Approach for Central Nervous System Tumor Classification from DNA Methylation

arXiv:2607. 01307v1 Announce Type: new Abstract: NA methylation profiling has become a powerful approach for central nervous system (CNS) tumor classification, yet important challenges remain regarding cross-cohort transferability, methodological correctness, and robust multiclass evaluation.

By Paulo R. Ferreira Jr., Lucas Coutinho Freitas, La\'is dos Santos Gon\c{c}alves, William Borges Domingues, Lucas Petitemberte de Souza, Mariana B. Michalowski, Vinicius F. Campos
arXiv AI
Aug 24

CellPath-Bench: A Multidimensional Benchmark for Whole-Slide Cellular Representations in Pathology Foundation Models

CellPath-Bench is a new benchmark that evaluates whole-slide cellular representations in pathology foundation models (PFMs) by using 25 spatially aligned H&E–Xenium tissue sections from 11 organs and over 7 million cells. It introduces metrics such as Cell Representation Advantage (CRA) and Cell Representation Transferability (CRT) to assess how well frozen PFMs encode cell-type information and generalize across tissue sections, datasets, and organs. The benchmark was applied to 30 PFMs, revealing significant model-dependent differences in cell-type decodability and cross-domain generalization, and offers a standardized framework for auditing cellular information in frozen PFM representations.

By Bokai Zhao, Yiyang Zhang, Hanqing Chao, Yawei Ma, Long Bai, Tai Ma, Minfeng Xu, Ming Song, Tianzi Jiang
arXiv AI
3d ago

CellMSA: Context Modeling for Single-Cell Representation Learning

CellMSA introduces a novel single‑cell representation learning framework that leverages a multiple‑sequence‑alignment‑inspired context model. For each target cell, it retrieves relevant cells across batches and related cell types, summarizing cross‑cell patterns into a context‑dependent gene‑pair representation that is fed into a pair‑aware encoder. Pretraining on a massive human single‑cell corpus (≈109 million cells) and subsequent benchmarks demonstrate consistent performance gains over existing methods.

By Suyuan Zhao, Minghao Liu, Yizhen Luo, Zaiqing Nie
arXiv Machine Learning
Jun 9

Integrating gene regulatory priors into Transformer attention with scTransformer for interpretable scRNA-seq analysis

arXiv:2606. 09558v1 Announce Type: cross Abstract: Motivation: Transformer-based models are increasingly applied to large-scale single-cell transcriptomics, showing strong performance through self-supervised learning on millions of cells.

By Mikele Milia, Louis Fabrice Tshimanga, Henning Mueller, Manfredo Atzori, Barbara Di Camillo
arXiv Machine Learning
Jun 19

Computational Methods and Challenges in Cell-Free DNA Analysis for Multi-Cancer Early Detection

arXiv:2606. 20174v1 Announce Type: new Abstract: Cell-free DNA (cfDNA) is a promising avenue for non-invasive multicancer early detection (MCED), in that, it can enable multiple cancer detection simultaneously from a single blood draw, with particular sensitivity to cancers that currently lack established screening programs.

By Nicko Starkey, Marcin W. Wojewodzic, Krzysztof Rzecki
arXiv Machine Learning
Sep 10

Multi-label versus multi-class classification of blood cells and their aggregates in microfluidic channels

arXiv:2609.07410v2 Announce Type: cross Abstract: Deformability cytometry (DC) is a type of imaging flow cytometry, which uses a camera-equipped device to measure cellular stiffness in addition to ot...

By Igor Zingman, Shada Abuhattum, Sara Kaliman, Maximilian Schl\"ogel, Paul M\"uller, Mark\'eta Kub\'ankov\'a, Nadine Str\"ohlein, Manuela Hauke, Lena Schn\"orer, Martin Kr\"ater, Jochen Guck
arXiv AI
3d ago

scTrilemma: Balancing Identity, Invariance, and Fidelity in Single-Cell Representation Learning

The paper introduces scTrilemma, a latent-bottleneck variational autoencoder designed to address the representation trilemma in single‑cell RNA‑seq data: preserving biological identity and state, remaining robust to nuisance context, and retaining gene‑level variation for expression analysis. scTrilemma routes expression‑derived variation to the embedding, decoder, or prior, gating gene tokens by expression and conditioning the prior on unlabeled pseudo‑bulk context, all under a single reconstruction objective without target annotations. In zero‑shot evaluations on successive CZ CELLxGENE Census releases, scTrilemma simultaneously satisfies all three demands, maintaining biological state, differential‑expression, and pathway structure across multiple disease settings, and latent interventions show context can be removed with minimal impact on other demands.

By Yunhak Oh, Yoonho Lee, Junseok Lee, Namkyeong Lee, Sang-Yeon Hwang, Yinhua Piao, Hyomin Kim, Seonghwan Kim, Jaechang Lim, Woo Youn Kim, Sungsoo Ahn, Chanyoung Park