arXiv Machine Learning By Jason X. Liu, Sebastian Ibarraran, Frank Hu, Soojung Yang, Xinyu A. Feng, Abigail Park, Anagha Aneesh, Lacramioara Bintu, Alexander R. Dunn, Grant M. Rotskoff

Generative modeling of intrinsically disordered protein regions by reinforcing sparse autoencoder features

Read the original on arXiv Machine Learning →

The paper introduces IDiom, an autoregressive protein language model trained on a large dataset of intrinsically disordered protein regions (IDRs) from AlphaFold, and demonstrates that it can generate sequences matching natural IDR composition, motifs, and disorder. It further presents RL‑SAE, a reinforcement learning approach that uses sparse autoencoder features to steer generation toward specific functional patterns, achieving high activation of targeted features and improved predicted subcellular localization and transcriptional activity. The combination of IDiom and RL‑SAE allows interpretable, composable IDR design by explicitly controlling function‑associated sequence features.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 28

Interpreting Latent Protein Language Model Features with Geometric Annotations

The paper introduces a scalable method to interpret sparse autoencoder (SAE) features in the ESM-2 protein language model by leveraging geometrically inspired features of the protein α‑carbon backbone. Across 8M layers of ESM-2, a false discovery rate–controlled analysis shows that local geometry is significantly associated with many SAE features, revealing substructure within known biological labels and enabling annotation of unannotated metagenomic proteins. Ablation experiments demonstrate that removing these geometric features shifts ESM-2’s predicted contact maps toward the descriptor, linking mechanistic interpretability with structural biology.

By Siddharth Setlur, Djordje Mihajlovic, Darrick Lee
arXiv Machine Learning
Sep 23

MT-ProtBERT: Multi-task Learning ProtBERT for Intrinsically Disordered Proteins Classification with Scarce Data

MT-ProtBERT is a multi‑task extension of ProtBERT designed for classifying intrinsically disordered proteins (IDPs) in low‑data settings. It combines Dynamic Window Masking, a Multi‑Scale 1D Convolutional classifier, and auxiliary biochemistry‑informed objectives to jointly optimize masked language modeling and domain‑specific tasks. In experiments on phosphorylation site prediction and protein compaction prediction, MT‑ProtBERT outperforms the RNN‑based IDP model PARROT across all limited‑data tasks.

By Jian Sun, Kingshuk Ghosh, Lilianna Houston, Mohammad H. Mahoor
arXiv Machine Learning
Aug 21

ProteinZero: Self-Improving Protein Generation via Online Reinforcement Learning

arXiv:2506. 07459v4 Announce Type: replace Abstract: Protein generative models have shown remarkable promise in protein design, yet their success rates remain constrained by reliance on curated sequence-structure datasets and by misalignment between supervised objectives and real design goals.

By Ziwen Wang, Jiajun Fan, Ruihan Guo, Thao Nguyen, Heng Ji, Ge Liu
arXiv Machine Learning
Sep 4

SimpleDesign: A Joint Model for Protein Sequence and Structure Codesign

SimpleDesign is a single-stage, end-to-end model for joint protein sequence and structure design that eliminates the need for multi-stage training. It combines discrete cross-entropy for sequences with a regression objective for structures, using a Mixture-of-Transformer architecture to handle modality-specific processing while maintaining global self-attention. Trained on over 2 million sequence-structure pairs, SimpleDesign achieves strong performance on co-design and unconditional generation benchmarks.

By Jiarui Lu, Yuyang Wang, Yizhe Zhang, Jiatao Gu, Navdeep Jaitly, Joshua M. Susskind, Miguel \'Angel Bautista
arXiv AI
Jul 23

Causal dictionary learning reveals and validates transcription-factor binding features in genomic language models

arXiv:2607. 19618v1 Announce Type: cross Abstract: Genomic language models achieve strong performance across regulatory-genomics tasks, yet what these models internally represent remains opaque, and the field lacks a principled procedure for verifying that an apparent ``concept'' inside a model is real rather than an artifact of sequence composition.

By Sarwan Ali