The paper introduces a scalable method to interpret sparse autoencoder (SAE) features in the ESM-2 protein language model by leveraging geometrically inspired features of the protein α‑carbon backbone. Across 8M layers of ESM-2, a false discovery rate–controlled analysis shows that local geometry is significantly associated with many SAE features, revealing substructure within known biological labels and enabling annotation of unannotated metagenomic proteins. Ablation experiments demonstrate that removing these geometric features shifts ESM-2’s predicted contact maps toward the descriptor, linking mechanistic interpretability with structural biology.
By Siddharth Setlur, Djordje Mihajlovic, Darrick Lee
MT-ProtBERT is a multi‑task extension of ProtBERT designed for classifying intrinsically disordered proteins (IDPs) in low‑data settings. It combines Dynamic Window Masking, a Multi‑Scale 1D Convolutional classifier, and auxiliary biochemistry‑informed objectives to jointly optimize masked language modeling and domain‑specific tasks. In experiments on phosphorylation site prediction and protein compaction prediction, MT‑ProtBERT outperforms the RNN‑based IDP model PARROT across all limited‑data tasks.
By Jian Sun, Kingshuk Ghosh, Lilianna Houston, Mohammad H. Mahoor
arXiv:2606. 16044v1 Announce Type: new Abstract: Protein language models (pLMs) can generate novel protein sequences with properties beyond those observed in nature, yet the mechanisms underlying protein generation remain poorly understood.
By Darin Tsui, William Deinzer, Daniel Saeedi, Amirali Aghazadeh
arXiv:2506. 07459v4 Announce Type: replace Abstract: Protein generative models have shown remarkable promise in protein design, yet their success rates remain constrained by reliance on curated sequence-structure datasets and by misalignment between supervised objectives and real design goals.
By Ziwen Wang, Jiajun Fan, Ruihan Guo, Thao Nguyen, Heng Ji, Ge Liu
SimpleDesign is a single-stage, end-to-end model for joint protein sequence and structure design that eliminates the need for multi-stage training. It combines discrete cross-entropy for sequences with a regression objective for structures, using a Mixture-of-Transformer architecture to handle modality-specific processing while maintaining global self-attention. Trained on over 2 million sequence-structure pairs, SimpleDesign achieves strong performance on co-design and unconditional generation benchmarks.
By Jiarui Lu, Yuyang Wang, Yizhe Zhang, Jiatao Gu, Navdeep Jaitly, Joshua M. Susskind, Miguel \'Angel Bautista
arXiv:2607. 19618v1 Announce Type: cross Abstract: Genomic language models achieve strong performance across regulatory-genomics tasks, yet what these models internally represent remains opaque, and the field lacks a principled procedure for verifying that an apparent ``concept'' inside a model is real rather than an artifact of sequence composition.
By Sarwan Ali