arXiv Machine Learning

MT-ProtBERT: Multi-task Learning ProtBERT for Intrinsically Disordered Proteins Classification with Scarce Data

MT-ProtBERT is a multi‑task extension of ProtBERT designed for classifying intrinsically disordered proteins (IDPs) in low‑data settings. It combines Dynamic Window Masking, a Multi‑Scale 1D Convolutional classifier, and auxiliary biochemistry‑informed objectives to jointly optimize masked language modeling and domain‑specific tasks. In experiments on phosphorylation site prediction and protein compaction prediction, MT‑ProtBERT outperforms the RNN‑based IDP model PARROT across all limited‑data tasks.

arXiv Machine Learning
1d ago

Generative modeling of intrinsically disordered protein regions by reinforcing sparse autoencoder features

The paper introduces IDiom, an autoregressive protein language model trained on a large dataset of intrinsically disordered protein regions (IDRs) from AlphaFold, and demonstrates that it can generate sequences matching natural IDR composition, motifs, and disorder. It further presents RL‑SAE, a reinforcement learning approach that uses sparse autoencoder features to steer generation toward specific functional patterns, achieving high activation of targeted features and improved predicted subcellular localization and transcriptional activity. The combination of IDiom and RL‑SAE allows interpretable, composable IDR design by explicitly controlling function‑associated sequence features.

By Jason X. Liu, Sebastian Ibarraran, Frank Hu, Soojung Yang, Xinyu A. Feng, Abigail Park, Anagha Aneesh, Lacramioara Bintu, Alexander R. Dunn, Grant M. Rotskoff
arXiv Machine Learning
Aug 28

Interpreting Latent Protein Language Model Features with Geometric Annotations

The paper introduces a scalable method to interpret sparse autoencoder (SAE) features in the ESM-2 protein language model by leveraging geometrically inspired features of the protein α‑carbon backbone. Across 8M layers of ESM-2, a false discovery rate–controlled analysis shows that local geometry is significantly associated with many SAE features, revealing substructure within known biological labels and enabling annotation of unannotated metagenomic proteins. Ablation experiments demonstrate that removing these geometric features shifts ESM-2’s predicted contact maps toward the descriptor, linking mechanistic interpretability with structural biology.

By Siddharth Setlur, Djordje Mihajlovic, Darrick Lee
arXiv Machine Learning
5d ago

GyroNovo: Error-Guided Fragment Imputation with Mass-Aware Attention for \textit{De Novo} Peptide Sequencing

GyroNovo is a new framework for de novo peptide sequencing that improves fragment imputation by guiding the process with decoder errors observed during training. It introduces mass-aware attention using rotary embeddings to encode pairwise mass differences between spectral peaks, and creates easy and hard augmented views of spectra to train the decoder under varying corruption levels. Experiments on NovoBench demonstrate significant gains, with about 9 percentage points higher peptide-level precision and 7 percentage points higher amino-acid-level precision compared to the state-of-the-art baseline.

By Abdellah El Mekki, Laks V. S. Lakshmanan, Muhammad Abdul-Mageed
Hugging Face Trending Papers
Sep 2

ProbeMatchDTI: Probe-Driven Multi-Scale Biochemical Pattern Matching for Drug-Target Interaction Prediction

ProbeMatchDTI introduces a probe-driven framework for drug‑target interaction prediction that preserves weak biochemical signals by using IterProbe to retain contextual states and BindingProbe to model cross‑entity complementarity at multiple scales. The method improves AUC‑ROC by 2.0% on BindingDB and 0.5% on DrugBank compared to prior biochemical representation learning approaches. Feature‑level analyses confirm the effectiveness of the probe-driven pattern matching, and the predictions are linked to an evidence‑guided downstream drug‑discovery workflow for candidate refinement and validation planning.