arXiv:2605.16581v2 Announce Type: replace
Abstract: Masked language modeling (MLM) is the standard objective for training protein language models, typically implemented by randomly masking individual...
By Thomas Walton, Ayan Goel, Amirali Aghazadeh
arXiv:2609.37675v1 Announce Type: new
Abstract: Protein Language Models (PLMs) have made remarkable progress following scaling laws established in natural language processing across sequence- and str...
By Biswajit Banerjee, Claudia Alvarez Carreno, Anton S. Petrov
The paper introduces a scalable method to interpret sparse autoencoder (SAE) features in the ESM-2 protein language model by leveraging geometrically inspired features of the protein α‑carbon backbone. Across 8M layers of ESM-2, a false discovery rate–controlled analysis shows that local geometry is significantly associated with many SAE features, revealing substructure within known biological labels and enabling annotation of unannotated metagenomic proteins. Ablation experiments demonstrate that removing these geometric features shifts ESM-2’s predicted contact maps toward the descriptor, linking mechanistic interpretability with structural biology.
By Siddharth Setlur, Djordje Mihajlovic, Darrick Lee
Protein language models (PLMs) have transferred the latest advances from natural language processing to computational biology. These models, trained on large corpora of protein sequence data, are widely used to translate amino acid sequences into latent-space embeddings, ready for use in diverse downstream tasks (DTs).
arXiv:2608. 12090v1 Announce Type: new Abstract: Protein language models (PLMs) have transferred the latest advances from natural language processing to computational biology.
By Roman Joeres, Ilya Senatorov, Olga V. Kalinina
The paper introduces a method that uses orthogonal projection to remove the influence of known biochemical features from protein language model (PLM) embeddings, allowing the authors to assess how much these features contribute to protein fitness predictions. By applying this technique to high‑order and interaction effects, they demonstrate that eliminating these interpretable features reduces downstream classifier performance, indicating that PLM embeddings encode patterns correlated with biochemical properties. The authors also show that these biochemical features explain a substantial portion of the variance in the classifier’s predictions, suggesting that PLM embeddings capture biologically relevant information.
By Paulo Yanez Sarmiento, Pia Francesca Rissom, Manuel Pfeuffer, Marco Simnacher, Jordan F. Safer, Sumaiya Iqbal, Henrike O. Heyne, Nadja Klein, Bernhard Y. Renard
The paper investigates how guided protein language models can collapse onto off‑manifold representations when heavily steered to optimize a property. This collapse causes generated sequences to become low‑complexity and statistically similar to random amino‑acid input, yet the property oracle may still rate them highly. The authors propose a cheap, training‑free Mahalanobis filtering step that removes such off‑manifold candidates, improving both property scores and structural plausibility without altering the generator.
ProtLingo is a protein language modeling framework that enhances a pretrained single‑sequence Transformer backbone with conditional local memory and sparse expert routing. It maps residue representations into discrete codes, composes local windows into latent N‑gram addresses, and retrieves reusable residual signals for recurring sequence contexts. The model also converts selected feed‑forward blocks into sparse Mixture‑of‑Experts layers, allowing residue‑dependent computation while activating only a subset of parameters, achieving competitive performance on protein fitness prediction, FLIP benchmarks, and supervised contact prediction with a 150M‑parameter backbone.
By Mingrui Li, Sixian Shen, Minzhang Li, Ruiyi Zhang, Kexin Zhang, Jiakai Zhang, Jingyi Yu
arXiv:2605. 16331v2 Announce Type: replace-cross Abstract: Protein language models are increasingly used to guide experimental and clinical decisions, yet it is often unclear whether a confident prediction reflects recognition of biological evidence or retrieval of a statistical default.
By Piotr Jedryszek, Oliver M. Crook
EvoLen is a tokenization method for DNA language models that incorporates evolutionary information to prioritize functional sequence patterns such as regulatory motifs. It groups DNA sequences by cross-species evolutionary signals, trains separate BPE tokenizers for each group, merges vocabularies with a rule that favors preserved patterns, and uses length-aware decoding with dynamic programming. Experiments show EvoLen better preserves functional motifs, differentiates genomic contexts, and aligns with evolutionary constraints while matching or surpassing standard BPE on various DNALM benchmarks.
By Nan Huang, Xiaoxiao Zhou, Junxia Cui, Mario Tapia-Pacheco, Tiffany Amariuta, Yang Li, Jingbo Shang
arXiv:2607. 12279v1 Announce Type: cross Abstract: Writing a sentence of exactly twelve words; ending a DNA sequence at the right codon; formatting an ASCII table.
By Jacob Dunefsky, Wes Gurnee, Emmanuel Ameisen
The paper investigates a problem in guided protein language models where strong guidance causes the model’s internal representations to collapse onto a region indistinguishable from random amino‑acid input, leading to low‑complexity sequences that still score well on the targeted property. The authors identify this off‑manifold collapse as a detectable signature and propose a post‑hoc filtering technique—Mahalanobis filtering—that removes atypical candidates based on a density prior over natural activations. This simple, training‑free step improves both property scores and structural plausibility across different guidance methods without altering the generator.
By Shuibai Zhang, Xinchi Liu, Fred Zhangzhi Peng, Zhihan Yang, Shutong Wu, Yingzi Ma, Jiawei Zhang