PFArena is a new benchmark for evaluating language models in protein modification tasks, featuring four controlled interfaces that span single‑mutant generation and multi‑mutant ranking. It incorporates varying levels of mutation fitness data to represent four research scenarios with different amounts of prior experimental context. The benchmark tests six protein language models, six large language models, and five LLM‑based agents, finding that PLMs excel at open‑ended single‑mutant generation while LLMs and agents perform best in multi‑mutant ranking when target‑specific data are available, yet all struggle as search space and mutation depth grow.
By Yawen Ouyang, Xinbo Zhang, Ziyuan Ma, Yixin Wu, Wenbin Liao, Feiran Zhang, Wenjie Li, Lihao Wang, Hao Wang, Xiaoqing Zheng, Xuefeng Yan, Lei Bai, Ya-Qin Zhang, Shuyi Zhang, Wei-Ying Ma, Dahua Lin, Bowen Zhou, Hao Zhou
Protein modification requires navigating an immense sequence space, yet wet-lab validation remains low-throughput and costly. Although computational paradigms including protein language models (PLMs),...
The paper introduces a method that uses orthogonal projection to remove the influence of known biochemical features from protein language model (PLM) embeddings, allowing the authors to assess how much these features contribute to protein fitness predictions. By applying this technique to high‑order and interaction effects, they demonstrate that eliminating these interpretable features reduces downstream classifier performance, indicating that PLM embeddings encode patterns correlated with biochemical properties. The authors also show that these biochemical features explain a substantial portion of the variance in the classifier’s predictions, suggesting that PLM embeddings capture biologically relevant information.
By Paulo Yanez Sarmiento, Pia Francesca Rissom, Manuel Pfeuffer, Marco Simnacher, Jordan F. Safer, Sumaiya Iqbal, Henrike O. Heyne, Nadja Klein, Bernhard Y. Renard
Q-BIOLAT is a framework that converts pretrained protein-language-model embeddings into compact binary codes and trains a quadratic unconstrained binary optimization (QUBO) surrogate with unary and pairwise latent interactions for protein fitness optimization. The study demonstrates that binary encodings with similar predictive accuracy can produce different Hamming neighborhoods, affecting local optima and search trajectories, and shows that PCA followed by per‑coordinate median thresholding yields a more balanced binary space than AE/VAE baselines. Experimental evaluation on GFP and AAV fitness landscapes from ProteinGym confirms that simulated annealing, genetic algorithms, and greedy hill climbing can retrieve high‑percentile variants, with decoded candidates reported via surrogate‑predicted scores.
By Truong-Son Hy
arXiv:2609.37808v1 Announce Type: new
Abstract: Protein optimization aims to discover high-fitness sequences under a limited experimental budget. Existing machine-learning methods use task-specific p...
By Zefeng Lin, Xianyong Fang, Tianfan Fu, Xiaohua Xu
arXiv:2605.16581v2 Announce Type: replace
Abstract: Masked language modeling (MLM) is the standard objective for training protein language models, typically implemented by randomly masking individual...
By Thomas Walton, Ayan Goel, Amirali Aghazadeh
arXiv:2606. 02386v1 Announce Type: new Abstract: Protein language models (PLMs) are passive oracles: they generate sequences in a single forward pass with no mechanism to consult external biophysical feedback or redirect generation when a candidate violates thermodynamic or structural constraints.
By Sahil Rahman, Maxx Richard Rahman
Protein language models (PLMs) have transferred the latest advances from natural language processing to computational biology. These models, trained on large corpora of protein sequence data, are widely used to translate amino acid sequences into latent-space embeddings, ready for use in diverse downstream tasks (DTs).
arXiv:2608. 12090v1 Announce Type: new Abstract: Protein language models (PLMs) have transferred the latest advances from natural language processing to computational biology.
By Roman Joeres, Ilya Senatorov, Olga V. Kalinina
arXiv:2606. 18961v1 Announce Type: new Abstract: Protein language models (PLMs) have emerged as powerful tools for controllable biomolecular design, yet their post-training adaptation typically relies on costly wet-lab validation or curated preference datasets.
By Lanqing Li, Shentong Mo, Yang Yu, Pheng-Ann Heng
The paper introduces StructEvo, a structure-aware reinforcement learning framework designed to improve protein directed evolution. By using a delta-structure fusion encoder to approximate mutant structure features and a hierarchical action network aligned with protein structure, the method navigates the vast mutation space more effectively. StructEvo outperforms existing machine learning-assisted directed evolution techniques by 9.2% and 16.3% on two benchmarks and uncovers an experimentally validated epistasis pattern in GFP, underscoring the value of structural guidance.
By Zikun Nie, Suyuan Zhao, Yizhen Luo, Siqi Fan, Zaiqing Nie
arXiv:2606. 08100v1 Announce Type: new Abstract: Multimodal $\Delta\Delta G$ predictors integrating protein language models with inverse-folding representations achieve strong in-distribution accuracy on the Megascale dataset but exhibit limited robustness on out-of-distribution (OOD) proteins, persistent forward-reverse bias on paired-mutation benchmarks, and under-representation of rare stabilizing mutations.
By A Shivram, Aneesh S. Chivukula, Manik Gupta, Sourav Chowdhury