Entropy, Disagreement, and the Limits of Foundation Models in Genomics
arXiv:2604. 04287v2 Announce Type: replace Abstract: Foundation models in genomics have shown mixed success compared to their counterparts in natural language processing.
arXiv:2602. 17162v3 Announce Type: replace Abstract: Genomic Foundation Models (GFMs) typically rely on Masked Language Modeling (MLM) or Next-Token Prediction (NTP) to learn the "Laws of Nature".
arXiv:2604. 04287v2 Announce Type: replace Abstract: Foundation models in genomics have shown mixed success compared to their counterparts in natural language processing.
arXiv:2609.36952v1 Announce Type: cross Abstract: Large language models (LLMs) excel at token-level generation but may learn undesirable abstract semantics and lack comprehensive perception. LLM-JEPA...
arXiv:2609.37675v1 Announce Type: new Abstract: Protein Language Models (PLMs) have made remarkable progress following scaling laws established in natural language processing across sequence- and str...
arXiv:2606. 05173v1 Announce Type: cross Abstract: Masked language modelling (MLM) has been the dominant pre-training objective for text encoders since BERT, yet it encourages representations that are strongly anchored to surface-form token identity rather than deeper semantic structure.
arXiv:2605.07938v2 Announce Type: replace Abstract: Single-cell representation learning (SCRL) from gene expression data offers a way to uncover the complex regulatory logic underlying cellular funct...
CellMSA introduces a novel single‑cell representation learning framework that leverages a multiple‑sequence‑alignment‑inspired context model. For each target cell, it retrieves relevant cells across batches and related cell types, summarizing cross‑cell patterns into a context‑dependent gene‑pair representation that is fed into a pair‑aware encoder. Pretraining on a massive human single‑cell corpus (≈109 million cells) and subsequent benchmarks demonstrate consistent performance gains over existing methods.
arXiv:2610.01942v1 Announce Type: new Abstract: Predicting the future evolution of a scene is a fundamental capability for world modeling. Recent work has shown that operating in the feature space of...
arXiv:2607. 29378v1 Announce Type: cross Abstract: Large language models (LLMs) generate text by auto-regressively sampling the next token.
The paper examines how new vocabulary tokens are added to language models for generative recommendation tasks. It shows that the common practice of initializing these tokens as the mean of existing embeddings collapses them into a degenerate subspace, hindering fine‑tuning. The authors propose Grounded Token Initialization (GTI), which places new tokens at semantically meaningful positions in the pretrained embedding space using linguistic supervision, and demonstrate that GTI outperforms mean initialization and other adaptation methods across several benchmarks.
arXiv:2608.29335v1 Announce Type: new Abstract: Latent generative models typically follow a two-stage pipeline, training a variational autoencoder for reconstruction and then a generative model on th...
arXiv:2606. 15521v1 Announce Type: cross Abstract: Tokenization introduces representational redundancy: under a fixed token vocabulary, every byte string admits many valid token encodings, or segmentations, that decode to the same surface string.
The paper introduces MOSAIC, a large adversarial benchmark for detecting AI-generated text, and presents NeuroStat, a new framework that combines token‑level probabilistic logits with deep semantic hidden states from a single language model. NeuroStat fuses these signals via Macro‑State Residual Modulation and uses orthogonal and contrastive losses to learn complementary representations. Experiments show that NeuroStat outperforms existing methods on MOSAIC, achieving superior robustness against adversarial attacks.