arXiv AI By Zaifei Yang, Samuel Ping-Man Choi, James Kwok

Enhancing Protein-Protein Interaction Prediction with Hierarchical Motif-based Multimodal Protein Embedding

Read the original on arXiv AI →

arXiv:2606. 02629v1 Announce Type: cross Abstract: Protein-protein interactions (PPIs) are essential for many biological processes.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Jun 29

PairSAE: Mechanistic Interpretability from Pair Representations in Protein Co-Folding

arXiv:2606. 27440v1 Announce Type: new Abstract: Foundation models for structural biology have achieved remarkable performance in predicting biomolecular structure and show promise for the design of proteins and small molecules.

By Giosue Migliorini, Aristofanis Rontogiannis, Grigori Guitchounts, Nicholas Franklin, Axel Elaldi, Olivia Viessmann
arXiv Machine Learning
Aug 28

Interpreting Latent Protein Language Model Features with Geometric Annotations

The paper introduces a scalable method to interpret sparse autoencoder (SAE) features in the ESM-2 protein language model by leveraging geometrically inspired features of the protein α‑carbon backbone. Across 8M layers of ESM-2, a false discovery rate–controlled analysis shows that local geometry is significantly associated with many SAE features, revealing substructure within known biological labels and enabling annotation of unannotated metagenomic proteins. Ablation experiments demonstrate that removing these geometric features shifts ESM-2’s predicted contact maps toward the descriptor, linking mechanistic interpretability with structural biology.

By Siddharth Setlur, Djordje Mihajlovic, Darrick Lee
Hugging Face Trending Papers
Sep 2

Subcellularly Resolved Single-Cell Embedding Learning with Transcriptomic data, Protein Structure and Localization Information

The paper introduces a multimodal framework that learns subcellularly resolved cell embeddings by integrating RNA expression profiles, protein sequence representations, and protein structural information. It uses a cross‑attention architecture to model interactions across distinct subcellular compartments, producing embeddings that capture both molecular expression patterns and functional protein properties. This approach is presented as the first to jointly incorporate transcriptomic data, protein sequences, and structural knowledge within a unified cross‑modal learning paradigm.