arXiv:2606. 18703v1 Announce Type: new Abstract: Pretrained biological language models expose per-token probability distributions through masked-token prediction, providing the likelihood interface central to sequence design, variant scoring, and mechanistic interpretation.
By Yanjun Shao, Yundi Chen, Yashvi Patel, Aurelien Pelissier, Mar\'ia Rodr\'iguez Mart\'inez
arXiv:2609.37808v1 Announce Type: new
Abstract: Protein optimization aims to discover high-fitness sequences under a limited experimental budget. Existing machine-learning methods use task-specific p...
By Zefeng Lin, Xianyong Fang, Tianfan Fu, Xiaohua Xu
The paper introduces a method that uses orthogonal projection to remove the influence of known biochemical features from protein language model (PLM) embeddings, allowing the authors to assess how much these features contribute to protein fitness predictions. By applying this technique to high‑order and interaction effects, they demonstrate that eliminating these interpretable features reduces downstream classifier performance, indicating that PLM embeddings encode patterns correlated with biochemical properties. The authors also show that these biochemical features explain a substantial portion of the variance in the classifier’s predictions, suggesting that PLM embeddings capture biologically relevant information.
By Paulo Yanez Sarmiento, Pia Francesca Rissom, Manuel Pfeuffer, Marco Simnacher, Jordan F. Safer, Sumaiya Iqbal, Henrike O. Heyne, Nadja Klein, Bernhard Y. Renard
The paper introduces ORBIT, a framework for probing higher‑order epistasis in protein representations. ORBIT validates Walsh‑based diagnostics on synthetic landscapes, then applies them to the GB1 fitness landscape, comparing several models including ridge regression, MLPs, and Residual Interaction Tokenization (RIT). While no architecture differences were found in overall prediction performance, RIT notably increased pairwise token‑level accessibility, and deeper MLPs improved higher‑order functional recovery, revealing representation‑level changes hidden by conventional metrics.
By Maryam Rahimimovassagh, Ivan Garibay, Niloofar Yousefi
arXiv:2606. 16044v1 Announce Type: new Abstract: Protein language models (pLMs) can generate novel protein sequences with properties beyond those observed in nature, yet the mechanisms underlying protein generation remain poorly understood.
By Darin Tsui, William Deinzer, Daniel Saeedi, Amirali Aghazadeh
arXiv:2609.24538v1 Announce Type: new
Abstract: Functional annotation of newly sequenced proteins remains a bottleneck in molecular biology: the number of sequences in public repositories grows far f...
By Demian Pavlyshenko, Bohdan Pavlyshenko
arXiv:2508. 20330v5 Announce Type: replace Abstract: Combinatorial optimization problems are ubiquitous in science and engineering.
By Zohair Shafi, Serdar Kadioglu
arXiv:2609.14418v1 Announce Type: cross
Abstract: Dynamic multi-mode resource-constrained project scheduling requires decisions to be made under precedence constraints, limited resources, multiple ex...
By Yuan Tian, Yi Mei, Mengjie Zhang
arXiv:2606. 08100v1 Announce Type: new Abstract: Multimodal $\Delta\Delta G$ predictors integrating protein language models with inverse-folding representations achieve strong in-distribution accuracy on the Megascale dataset but exhibit limited robustness on out-of-distribution (OOD) proteins, persistent forward-reverse bias on paired-mutation benchmarks, and under-representation of rare stabilizing mutations.
By A Shivram, Aneesh S. Chivukula, Manik Gupta, Sourav Chowdhury
arXiv:2606. 02386v1 Announce Type: new Abstract: Protein language models (PLMs) are passive oracles: they generate sequences in a single forward pass with no mechanism to consult external biophysical feedback or redirect generation when a candidate violates thermodynamic or structural constraints.
By Sahil Rahman, Maxx Richard Rahman
arXiv:2606. 18961v1 Announce Type: new Abstract: Protein language models (PLMs) have emerged as powerful tools for controllable biomolecular design, yet their post-training adaptation typically relies on costly wet-lab validation or curated preference datasets.
By Lanqing Li, Shentong Mo, Yang Yu, Pheng-Ann Heng
arXiv:2606. 05139v1 Announce Type: new Abstract: The rapid advancement of high-throughput sequencing has led to large, high-dimensional omics datasets.
By Luca Thale-Bombien, Jan Ewald, Ralf K\"onig, Aaron Klein