arXiv:2604. 04287v2 Announce Type: replace Abstract: Foundation models in genomics have shown mixed success compared to their counterparts in natural language processing.
By Maxime Rochkoulets, Lovro Vr\v{c}ek, Mile \v{S}iki\'c
arXiv:2602. 17162v3 Announce Type: replace Abstract: Genomic Foundation Models (GFMs) typically rely on Masked Language Modeling (MLM) or Next-Token Prediction (NTP) to learn the "Laws of Nature".
By Ariel Larey, Elay Dahan, Amit Bleiweiss, Raizy Kellerman, Guy Leib, Omri Nayshool, Dan Ofer, Tal Zinger, Dan Dominissini, Gideon Rechavi, Nicole Bussola, Simon Lee, Shane O'Connell, Dung Hoang, Marissa Wirth, Alexander W. Charney, Nati Daniel, Yoli Shavit
The paper proposes a three‑stage training pipeline that begins with procedural pretraining on abstract, procedurally generated data, followed by molecular pretraining on SMILES, and finally downstream fine‑tuning for molecular property prediction. Experiments show that procedural pretraining improves downstream performance—e.g., a 4.8% error reduction on Lipophilicity—especially when labeled data are scarce, and that the benefit peaks at an intermediate procedural training budget. Analysis indicates that transferable knowledge resides mainly in attention layers, while feed‑forward layers may over‑specialize.
By Moritz Friedemann, Zachary Shinnick, Philip Torr, Bruno Andreis
arXiv:2603. 25062v2 Announce Type: replace Abstract: Autoregressive molecular models assign probability to molecular serializations even though chemical identity is invariant to serialization.
By Xinyu Wang, Fei Dou, Jinbo Bi, Minghu Song
arXiv:2510. 14217v2 Announce Type: replace Abstract: The spectral properties of feature embeddings offer critical insights into model generalization and representation quality.
By Asma Jamali, Tin Sum Cheng, Rodrigo A. Vargas-Hern\'andez
WEECFP-SuRGE introduces a position‑aware substructure encoding method that combines tokenized hierarchical Morgan fingerprints with graph‑distance‑dependent rotations applied at the input and within transformer self‑attention. The approach captures local chemistry, long‑range interactions, and molecular topology without requiring external pretraining or 3‑D conformer generation. Benchmarks on MoleculeNet and the Therapeutic Data Commons ADMET datasets show competitive performance, and a reconstruction procedure correctly identifies constitutional isomers for 92.6% of a 4,200‑molecule library.
By Robert Epps