Linear Fitness Subspace in Protein Language Models Enables Sample-Efficient Directed Evolution
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
PFArena is a new benchmark for evaluating language models in protein modification tasks, featuring four controlled interfaces that span single‑mutant generation and multi‑mutant ranking. It incorporates varying levels of mutation fitness data to represent four research scenarios with different amounts of prior experimental context. The benchmark tests six protein language models, six large language models, and five LLM‑based agents, finding that PLMs excel at open‑ended single‑mutant generation while LLMs and agents perform best in multi‑mutant ranking when target‑specific data are available, yet all struggle as search space and mutation depth grow.
Protein modification requires navigating an immense sequence space, yet wet-lab validation remains low-throughput and costly. Although computational paradigms including protein language models (PLMs),...
The paper introduces a method that uses orthogonal projection to remove the influence of known biochemical features from protein language model (PLM) embeddings, allowing the authors to assess how much these features contribute to protein fitness predictions. By applying this technique to high‑order and interaction effects, they demonstrate that eliminating these interpretable features reduces downstream classifier performance, indicating that PLM embeddings encode patterns correlated with biochemical properties. The authors also show that these biochemical features explain a substantial portion of the variance in the classifier’s predictions, suggesting that PLM embeddings capture biologically relevant information.
Q-BIOLAT is a framework that converts pretrained protein-language-model embeddings into compact binary codes and trains a quadratic unconstrained binary optimization (QUBO) surrogate with unary and pairwise latent interactions for protein fitness optimization. The study demonstrates that binary encodings with similar predictive accuracy can produce different Hamming neighborhoods, affecting local optima and search trajectories, and shows that PCA followed by per‑coordinate median thresholding yields a more balanced binary space than AE/VAE baselines. Experimental evaluation on GFP and AAV fitness landscapes from ProteinGym confirms that simulated annealing, genetic algorithms, and greedy hill climbing can retrieve high‑percentile variants, with decoded candidates reported via surrogate‑predicted scores.
arXiv:2609.37808v1 Announce Type: new Abstract: Protein optimization aims to discover high-fitness sequences under a limited experimental budget. Existing machine-learning methods use task-specific p...
arXiv:2605.16581v2 Announce Type: replace Abstract: Masked language modeling (MLM) is the standard objective for training protein language models, typically implemented by randomly masking individual...