arXiv:2609.37675v1 Announce Type: new
Abstract: Protein Language Models (PLMs) have made remarkable progress following scaling laws established in natural language processing across sequence- and str...
By Biswajit Banerjee, Claudia Alvarez Carreno, Anton S. Petrov
ProtLingo is a protein language modeling framework that enhances a pretrained single‑sequence Transformer backbone with conditional local memory and sparse expert routing. It maps residue representations into discrete codes, composes local windows into latent N‑gram addresses, and retrieves reusable residual signals for recurring sequence contexts. The model also converts selected feed‑forward blocks into sparse Mixture‑of‑Experts layers, allowing residue‑dependent computation while activating only a subset of parameters, achieving competitive performance on protein fitness prediction, FLIP benchmarks, and supervised contact prediction with a 150M‑parameter backbone.
By Mingrui Li, Sixian Shen, Minzhang Li, Ruiyi Zhang, Kexin Zhang, Jiakai Zhang, Jingyi Yu
arXiv:2606. 02629v1 Announce Type: cross Abstract: Protein-protein interactions (PPIs) are essential for many biological processes.
By Zaifei Yang, Samuel Ping-Man Choi, James Kwok
SimpleDesign is a single-stage, end-to-end model for joint protein sequence and structure design that eliminates the need for multi-stage training. It combines discrete cross-entropy for sequences with a regression objective for structures, using a Mixture-of-Transformer architecture to handle modality-specific processing while maintaining global self-attention. Trained on over 2 million sequence-structure pairs, SimpleDesign achieves strong performance on co-design and unconditional generation benchmarks.
By Jiarui Lu, Yuyang Wang, Yizhe Zhang, Jiatao Gu, Navdeep Jaitly, Joshua M. Susskind, Miguel \'Angel Bautista
arXiv:2608. 12090v1 Announce Type: new Abstract: Protein language models (PLMs) have transferred the latest advances from natural language processing to computational biology.
By Roman Joeres, Ilya Senatorov, Olga V. Kalinina
arXiv:2608. 16094v1 Announce Type: new Abstract: Accurate protein structure prediction is fundamental to structural biology because protein structure underlies molecular function and provides a basis for mechanistic interpretation.
By Wengan He, Yongsheng Luo, Lihong Jiang, Wenhui Xu, Yu Li
Protein language models (PLMs) have transferred the latest advances from natural language processing to computational biology. These models, trained on large corpora of protein sequence data, are widely used to translate amino acid sequences into latent-space embeddings, ready for use in diverse downstream tasks (DTs).
The study evaluates 4‑bit quantization and low‑rank adapter fine‑tuning (QLoRA) on several large protein language models, finding that many model‑task pairs retain over 90% of full fine‑tuning performance while achieving up to 90% GPU memory savings. QLoRA preserves early‑layer representations and induces task‑specific changes in later layers, closely resembling full fine‑tuning with smaller representational shifts. For generative models, 4‑bit quantization largely maintains structural and sequence‑level properties, though token‑level analysis reveals model‑dependent changes in autoregressive output distributions.
By Ilan Yaniv Zeisler, Sebastian Clancy, Pouriya Bayat, Saaim Raad, Ivan Kraskov, Matthew Xie, Vivian White, Spencer Perkins, Serena Singh, Sepehr Bayat, Keith Pardee
arXiv:2605. 00182v3 Announce Type: replace Abstract: Proteins are shaped by gradual evolution under biophysical and functional constraints.
By Xinyou Wang, Liang Hong, Jiasheng Ye, Zaixiang Zheng, Yu Li, Shujian Huang, Quanquan Gu
SymFold introduces a symmetric dual‑path architecture that combines protein language models (PLMs) and multimodal protein language models (MPLMs) to iteratively guide protein sequence generation for inverse folding. By leveraging pretrained sequence evolution knowledge from PLMs and structural knowledge from MPLMs, the method improves upon the traditional serial pipeline where structure encoders produce coarse sequences refined by PLMs. Experiments on standard inverse‑folding benchmarks show state‑of‑the‑art performance, and ablation studies confirm the effectiveness of the symmetric design.
By Handong Wang, Jiaxin Qi, Baisheng Lai, Jianqiang Huang
arXiv:2511. 09465v4 Announce Type: replace-cross Abstract: Diffusion and flow matching approaches to generative modeling have shown promise in domains where the state space is continuous, such as image generation or protein folding & design, and discrete, exemplified by diffusion large language models.
By Lukas Billera, Hedwig Nora Nordlinder, Jack Collier Ryder, Anton Oresten, Aron St{\aa}lmarck, Theodor Mosetti Bj\"ork, Ben Murrell
arXiv:2607. 22777v1 Announce Type: cross Abstract: Protein language models learn transferable sequence representations.
By Chen Wang, Boming Kang, Qinghua Cui