arXiv Machine Learning By Dmytro Rizdvanetskyi, Nathan Ross, Pavlo Lutsik

Data-Driven Soft Labeling Scales DNA Read Classification to Whole-Body Cell-Type Deconvolution

Read the original on arXiv Machine Learning →

arXiv:2607. 04987v1 Announce Type: new Abstract: Cell-type deconvolution, the task of estimating the proportions of constituent cell types in a heterogeneous biological sample, is a core problem in computational biology.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jul 3

A Novel Machine Learning Approach for Central Nervous System Tumor Classification from DNA Methylation

arXiv:2607. 01307v1 Announce Type: new Abstract: NA methylation profiling has become a powerful approach for central nervous system (CNS) tumor classification, yet important challenges remain regarding cross-cohort transferability, methodological correctness, and robust multiclass evaluation.

By Paulo R. Ferreira Jr., Lucas Coutinho Freitas, La\'is dos Santos Gon\c{c}alves, William Borges Domingues, Lucas Petitemberte de Souza, Mariana B. Michalowski, Vinicius F. Campos
arXiv AI
Aug 24

CellPath-Bench: A Multidimensional Benchmark for Whole-Slide Cellular Representations in Pathology Foundation Models

CellPath-Bench is a new benchmark that evaluates whole-slide cellular representations in pathology foundation models (PFMs) by using 25 spatially aligned H&E–Xenium tissue sections from 11 organs and over 7 million cells. It introduces metrics such as Cell Representation Advantage (CRA) and Cell Representation Transferability (CRT) to assess how well frozen PFMs encode cell-type information and generalize across tissue sections, datasets, and organs. The benchmark was applied to 30 PFMs, revealing significant model-dependent differences in cell-type decodability and cross-domain generalization, and offers a standardized framework for auditing cellular information in frozen PFM representations.

By Bokai Zhao, Yiyang Zhang, Hanqing Chao, Yawei Ma, Long Bai, Tai Ma, Minfeng Xu, Ming Song, Tianzi Jiang
arXiv AI
3d ago

CellMSA: Context Modeling for Single-Cell Representation Learning

CellMSA introduces a novel single‑cell representation learning framework that leverages a multiple‑sequence‑alignment‑inspired context model. For each target cell, it retrieves relevant cells across batches and related cell types, summarizing cross‑cell patterns into a context‑dependent gene‑pair representation that is fed into a pair‑aware encoder. Pretraining on a massive human single‑cell corpus (≈109 million cells) and subsequent benchmarks demonstrate consistent performance gains over existing methods.

By Suyuan Zhao, Minghao Liu, Yizhen Luo, Zaiqing Nie