arXiv Machine Learning By Maxime Rochkoulets, Lovro Vr\v{c}ek, Mile \v{S}iki\'c

Entropy, Disagreement, and the Limits of Foundation Models in Genomics

Read the original on arXiv Machine Learning →

arXiv:2604. 04287v2 Announce Type: replace Abstract: Foundation models in genomics have shown mixed success compared to their counterparts in natural language processing.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Aug 12

JEPA-DNA: Grounding Genomic Foundation Models through Joint-Embedding Predictive Architectures

arXiv:2602. 17162v3 Announce Type: replace Abstract: Genomic Foundation Models (GFMs) typically rely on Masked Language Modeling (MLM) or Next-Token Prediction (NTP) to learn the "Laws of Nature".

By Ariel Larey, Elay Dahan, Amit Bleiweiss, Raizy Kellerman, Guy Leib, Omri Nayshool, Dan Ofer, Tal Zinger, Dan Dominissini, Gideon Rechavi, Nicole Bussola, Simon Lee, Shane O'Connell, Dung Hoang, Marissa Wirth, Alexander W. Charney, Nati Daniel, Yoli Shavit
arXiv Machine Learning
Sep 1

Learning Representations through Token Prediction: Geometry, Approximation, and Downstream Guarantees

The paper investigates why token prediction, a common pre‑training objective for language models, yields useful representations. It introduces a statistical framework linking token prediction accuracy to the geometry of token embeddings, showing that accurate predictions organize embeddings according to Hellinger distances between context distributions. The authors also propose a self‑consistency principle that refines contextual representations through repeated application of a shared block, and provide downstream guarantees for token generation, community recovery, and linear classification.

By Shulei Wang
arXiv Statistics ML
2d ago

VANDAM: Viewing a nucleotide sequence with DNA molecular priors

VANDAM is a framework that augments Genomic Foundation Models by incorporating DNA molecular priors into self‑supervised training. It predicts regional molecular properties from pooled representations and, when functional labels are available, injects local features at the input. The approach consistently improves downstream performance across multiple architecture families and genomic tasks, and probing experiments show that the priors generalize to unseen molecular properties.

By Jeremy Levy, Ariel Larey, Yury Nahshan, Raizy Kellerman, Elay Dahan, Amit Bleiweiss, Guy Leib, Omri Nayshool, Dan Ofer, Tal Zinger, Dan Dominissini, Gideon Rechavi, Marissa Wirth, Simon Lee, Dung Hoang, Noam D. Beckmann, Shane O'Connell, Nicole Bussola, Alexander W. Charney, Yoli Shavit, Nati Daniel