arXiv AI By SiYuan Ma, Canran Xiao, Zikai Xiao, Albert Gao, Liang He, Xuan-Yu Wang, Shuying Cao, Xiaojun Jia

Linear Fitness Subspace in Protein Language Models Enables Sample-Efficient Directed Evolution

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv AI
Sep 25

PFArena: Benchmarking Language Models for Protein Modification

PFArena is a new benchmark for evaluating language models in protein modification tasks, featuring four controlled interfaces that span single‑mutant generation and multi‑mutant ranking. It incorporates varying levels of mutation fitness data to represent four research scenarios with different amounts of prior experimental context. The benchmark tests six protein language models, six large language models, and five LLM‑based agents, finding that PLMs excel at open‑ended single‑mutant generation while LLMs and agents perform best in multi‑mutant ranking when target‑specific data are available, yet all struggle as search space and mutation depth grow.

By Yawen Ouyang, Xinbo Zhang, Ziyuan Ma, Yixin Wu, Wenbin Liao, Feiran Zhang, Wenjie Li, Lihao Wang, Hao Wang, Xiaoqing Zheng, Xuefeng Yan, Lei Bai, Ya-Qin Zhang, Shuyi Zhang, Wei-Ying Ma, Dahua Lin, Bowen Zhou, Hao Zhou
arXiv Machine Learning
Aug 27

Interpreting Protein Language Model Embeddings via Orthogonal Projection for Protein Fitness Prediction

The paper introduces a method that uses orthogonal projection to remove the influence of known biochemical features from protein language model (PLM) embeddings, allowing the authors to assess how much these features contribute to protein fitness predictions. By applying this technique to high‑order and interaction effects, they demonstrate that eliminating these interpretable features reduces downstream classifier performance, indicating that PLM embeddings encode patterns correlated with biochemical properties. The authors also show that these biochemical features explain a substantial portion of the variance in the classifier’s predictions, suggesting that PLM embeddings capture biologically relevant information.

By Paulo Yanez Sarmiento, Pia Francesca Rissom, Manuel Pfeuffer, Marco Simnacher, Jordan F. Safer, Sumaiya Iqbal, Henrike O. Heyne, Nadja Klein, Bernhard Y. Renard
arXiv Machine Learning
Sep 17

Q-BIOLAT: Binary Latent Protein Fitness Landscapes for QUBO-Based Optimization

Q-BIOLAT is a framework that converts pretrained protein-language-model embeddings into compact binary codes and trains a quadratic unconstrained binary optimization (QUBO) surrogate with unary and pairwise latent interactions for protein fitness optimization. The study demonstrates that binary encodings with similar predictive accuracy can produce different Hamming neighborhoods, affecting local optima and search trajectories, and shows that PCA followed by per‑coordinate median thresholding yields a more balanced binary space than AE/VAE baselines. Experimental evaluation on GFP and AAV fitness landscapes from ProteinGym confirms that simulated annealing, genetic algorithms, and greedy hill climbing can retrieve high‑percentile variants, with decoded candidates reported via surrogate‑predicted scores.

By Truong-Son Hy