Task- and dataset-specific information in protein language models
arXiv:2608. 12090v1 Announce Type: new Abstract: Protein language models (PLMs) have transferred the latest advances from natural language processing to computational biology.
Protein language models (PLMs) have transferred the latest advances from natural language processing to computational biology. These models, trained on large corpora of protein sequence data, are widely used to translate amino acid sequences into latent-space embeddings, ready for use in diverse downstream tasks (DTs).
arXiv:2608. 12090v1 Announce Type: new Abstract: Protein language models (PLMs) have transferred the latest advances from natural language processing to computational biology.
The paper introduces a method that uses orthogonal projection to remove the influence of known biochemical features from protein language model (PLM) embeddings, allowing the authors to assess how much these features contribute to protein fitness predictions. By applying this technique to high‑order and interaction effects, they demonstrate that eliminating these interpretable features reduces downstream classifier performance, indicating that PLM embeddings encode patterns correlated with biochemical properties. The authors also show that these biochemical features explain a substantial portion of the variance in the classifier’s predictions, suggesting that PLM embeddings capture biologically relevant information.
The paper introduces Murmur2Vec, a lightweight, alignment‑free embedding that uses k‑mer counts hashed with MurmurHash to create a compact representation for biological sequences. It provides a full theoretical analysis, including bias/variance formulas, a Johnson–Lindenstrauss‑style concentration bound, and an excess‑risk bound that clarifies the trade‑off between hash‑table size and classifier performance. Empirically, Murmur2Vec matches or surpasses a fine‑tuned 650M‑parameter ESM‑2 protein language model across several classification tasks, including SARS‑CoV‑2 spike lineage and HIV‑1 Env subtype identification.
The paper introduces a scalable method to interpret sparse autoencoder (SAE) features in the ESM-2 protein language model by leveraging geometrically inspired features of the protein α‑carbon backbone. Across 8M layers of ESM-2, a false discovery rate–controlled analysis shows that local geometry is significantly associated with many SAE features, revealing substructure within known biological labels and enabling annotation of unannotated metagenomic proteins. Ablation experiments demonstrate that removing these geometric features shifts ESM-2’s predicted contact maps toward the descriptor, linking mechanistic interpretability with structural biology.
The paper investigates how to effectively pre‑train language models when the data budget is limited but compute is plentiful. It shows that increasing model size only improves performance up to an optimal point, after which overfitting degrades generalization, and that this optimal size varies with both the data budget and downstream tasks. To overcome the inefficiencies of standard Transformers in this regime, the authors propose recursive Transformers that reuse a shared block across depth and employ factorized embeddings, achieving better results than standard models on 10M–100M word pre‑training budgets and competitive performance with BabyLM Challenge 2025 winners.
arXiv:2609.37675v1 Announce Type: new Abstract: Protein Language Models (PLMs) have made remarkable progress following scaling laws established in natural language processing across sequence- and str...
arXiv:2512. 15133v3 Announce Type: replace-cross Abstract: Proteins inherently possess a consistent sequence-structure duality.
arXiv:2603.18908v5 Announce Type: replace Abstract: Independently trained language models often learn compatible late-stage representations, despite differences in training objectives, architectures,...
arXiv:2607. 22777v1 Announce Type: cross Abstract: Protein language models learn transferable sequence representations.
arXiv:2608.30315v1 Announce Type: new Abstract: Token embeddings are the basic representational units that connect discrete tokens with continuous computation in language models. Although modern lang...
arXiv:2608. 06111v1 Announce Type: cross Abstract: Positional embeddings (PE) in Transformers encode token distance and order but are largely agnostic to \textit{syntactic structure}.
arXiv:2604. 26170v2 Announce Type: replace Abstract: Adapting large language models (LLMs) to a targeted task efficiently and effectively remains a fundamental challenge.