arXiv:2602. 17162v3 Announce Type: replace Abstract: Genomic Foundation Models (GFMs) typically rely on Masked Language Modeling (MLM) or Next-Token Prediction (NTP) to learn the "Laws of Nature".
By Ariel Larey, Elay Dahan, Amit Bleiweiss, Raizy Kellerman, Guy Leib, Omri Nayshool, Dan Ofer, Tal Zinger, Dan Dominissini, Gideon Rechavi, Nicole Bussola, Simon Lee, Shane O'Connell, Dung Hoang, Marissa Wirth, Alexander W. Charney, Nati Daniel, Yoli Shavit
The paper investigates why token prediction, a common pre‑training objective for language models, yields useful representations. It introduces a statistical framework linking token prediction accuracy to the geometry of token embeddings, showing that accurate predictions organize embeddings according to Hellinger distances between context distributions. The authors also propose a self‑consistency principle that refines contextual representations through repeated application of a shared block, and provide downstream guarantees for token generation, community recovery, and linear classification.
By Shulei Wang
VANDAM is a framework that augments Genomic Foundation Models by incorporating DNA molecular priors into self‑supervised training. It predicts regional molecular properties from pooled representations and, when functional labels are available, injects local features at the input. The approach consistently improves downstream performance across multiple architecture families and genomic tasks, and probing experiments show that the priors generalize to unseen molecular properties.
By Jeremy Levy, Ariel Larey, Yury Nahshan, Raizy Kellerman, Elay Dahan, Amit Bleiweiss, Guy Leib, Omri Nayshool, Dan Ofer, Tal Zinger, Dan Dominissini, Gideon Rechavi, Marissa Wirth, Simon Lee, Dung Hoang, Noam D. Beckmann, Shane O'Connell, Nicole Bussola, Alexander W. Charney, Yoli Shavit, Nati Daniel
arXiv:2608. 12090v1 Announce Type: new Abstract: Protein language models (PLMs) have transferred the latest advances from natural language processing to computational biology.
By Roman Joeres, Ilya Senatorov, Olga V. Kalinina
arXiv:2608.23551v1 Announce Type: cross
Abstract: Recent advances in continuous diffusion and flow-based language models (LMs) have achieved performance competitive with discrete LMs. However, existi...
By Na Li, Yuchen Jiao, Changxiao Cai, Gen Li
Protein language models (PLMs) have transferred the latest advances from natural language processing to computational biology. These models, trained on large corpora of protein sequence data, are widely used to translate amino acid sequences into latent-space embeddings, ready for use in diverse downstream tasks (DTs).
arXiv:2605.07938v2 Announce Type: replace
Abstract: Single-cell representation learning (SCRL) from gene expression data offers a way to uncover the complex regulatory logic underlying cellular funct...
By Sachini Weerasekara, Natasha Darras, Sagar Kamarthi, Colles Price, Jacqueline Isaacs
arXiv:2509.20702v3 Announce Type: replace-cross
Abstract: Recent advances in large language model (LLM) embeddings have enabled powerful representations for biological data, but most applications to...
By Hongqian Niu, Jordan Bryan, Jacob Williams, Hufeng Zhou, Zhun Deng, Haoyu Zhang, Xihao Li, Didong Li
arXiv:2608.30315v1 Announce Type: new
Abstract: Token embeddings are the basic representational units that connect discrete tokens with continuous computation in language models. Although modern lang...
By Junjie Yao, Liangkai Hang, Zhi-Qin John Xu
arXiv:2607. 04733v1 Announce Type: cross Abstract: Supervised fine-tuning (SFT) is the standard approach for adapting pretrained language models to downstream domains, yet it often improves target-domain behavior at the cost of degrading pre-existing capabilities.
By Yueyang Wang, Baolong Bi, Shuo Lu, Jingyuan Zhang
arXiv:2607. 08803v1 Announce Type: cross Abstract: The push toward large language models for biology (BioLM) has created a need for training corpora that can endow models with a genuine understanding of biology.
By Hyunjin Seo, Hyeon Hwang, Gyubok Lee, Jay Shin, Jimin Park, Taesoo Kim, Sanghoon Lee, Hongjoon Ahn, Sungjun Han, Sangwon Jung
CellMSA introduces a novel single‑cell representation learning framework that leverages a multiple‑sequence‑alignment‑inspired context model. For each target cell, it retrieves relevant cells across batches and related cell types, summarizing cross‑cell patterns into a context‑dependent gene‑pair representation that is fed into a pair‑aware encoder. Pretraining on a massive human single‑cell corpus (≈109 million cells) and subsequent benchmarks demonstrate consistent performance gains over existing methods.
By Suyuan Zhao, Minghao Liu, Yizhen Luo, Zaiqing Nie