arXiv Machine Learning By Calvin McCarter, Nick Bhattacharya, Sebastian W. Ober, Hunter Elliott

How to make the most of your masked language model for protein engineering

Read the original on arXiv Machine Learning →

The paper introduces a flexible sampling technique for masked language models (MLMs) called stochastic beam search, which leverages MLMs’ efficiency in evaluating the pseudo‑perplexity of a sequence’s 1‑edit neighborhood. This method reframes generation as whole‑sequence evaluation, allowing guidance across multiple optimization objectives. Extensive in‑silico and in‑vitro tests on antibody therapeutics demonstrate that the sampling strategy significantly influences outcomes, highlighting the need for further research in this area.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 25

PFArena: Benchmarking Language Models for Protein Modification

PFArena is a new benchmark for evaluating language models in protein modification tasks, featuring four controlled interfaces that span single‑mutant generation and multi‑mutant ranking. It incorporates varying levels of mutation fitness data to represent four research scenarios with different amounts of prior experimental context. The benchmark tests six protein language models, six large language models, and five LLM‑based agents, finding that PLMs excel at open‑ended single‑mutant generation while LLMs and agents perform best in multi‑mutant ranking when target‑specific data are available, yet all struggle as search space and mutation depth grow.

By Yawen Ouyang, Xinbo Zhang, Ziyuan Ma, Yixin Wu, Wenbin Liao, Feiran Zhang, Wenjie Li, Lihao Wang, Hao Wang, Xiaoqing Zheng, Xuefeng Yan, Lei Bai, Ya-Qin Zhang, Shuyi Zhang, Wei-Ying Ma, Dahua Lin, Bowen Zhou, Hao Zhou
Hugging Face Trending Papers
Aug 19

Off-Manifold Collapse in Guided Protein Language Models

The paper investigates how guided protein language models can collapse onto off‑manifold representations when heavily steered to optimize a property. This collapse causes generated sequences to become low‑complexity and statistically similar to random amino‑acid input, yet the property oracle may still rate them highly. The authors propose a cheap, training‑free Mahalanobis filtering step that removes such off‑manifold candidates, improving both property scores and structural plausibility without altering the generator.