arXiv:2607. 01307v1 Announce Type: new Abstract: NA methylation profiling has become a powerful approach for central nervous system (CNS) tumor classification, yet important challenges remain regarding cross-cohort transferability, methodological correctness, and robust multiclass evaluation.
By Paulo R. Ferreira Jr., Lucas Coutinho Freitas, La\'is dos Santos Gon\c{c}alves, William Borges Domingues, Lucas Petitemberte de Souza, Mariana B. Michalowski, Vinicius F. Campos
arXiv:2607. 19426v2 Announce Type: replace-cross Abstract: Large single-cell datasets are expensive to store, curate, and repeatedly reuse for model training.
By Yaodi Luo, Peize He, Lingbei Meng, Bowen Han, Zheng Lu, Jianqing Zhu, Lian Zhang
arXiv:2607. 19426v1 Announce Type: cross Abstract: Single-cell datasets are increasingly costly to store, audit, and reuse for model training.
By Yaodi Luo, Peize He, Bowen Han, Lingbei Mengg
CellPath-Bench is a new benchmark that evaluates whole-slide cellular representations in pathology foundation models (PFMs) by using 25 spatially aligned H&E–Xenium tissue sections from 11 organs and over 7 million cells. It introduces metrics such as Cell Representation Advantage (CRA) and Cell Representation Transferability (CRT) to assess how well frozen PFMs encode cell-type information and generalize across tissue sections, datasets, and organs. The benchmark was applied to 30 PFMs, revealing significant model-dependent differences in cell-type decodability and cross-domain generalization, and offers a standardized framework for auditing cellular information in frozen PFM representations.
By Bokai Zhao, Yiyang Zhang, Hanqing Chao, Yawei Ma, Long Bai, Tai Ma, Minfeng Xu, Ming Song, Tianzi Jiang
arXiv:2602. 15253v2 Announce Type: replace Abstract: Neural scaling laws -- power-law relationships between loss, model size, and data -- have been extensively documented for language and vision transformers, yet their existence in single-cell genomics remains largely unexplored.
By Ihor Kendiukhov
CellMSA introduces a novel single‑cell representation learning framework that leverages a multiple‑sequence‑alignment‑inspired context model. For each target cell, it retrieves relevant cells across batches and related cell types, summarizing cross‑cell patterns into a context‑dependent gene‑pair representation that is fed into a pair‑aware encoder. Pretraining on a massive human single‑cell corpus (≈109 million cells) and subsequent benchmarks demonstrate consistent performance gains over existing methods.
By Suyuan Zhao, Minghao Liu, Yizhen Luo, Zaiqing Nie