Towards Data Science

Proteins: A Mosaic Pattern to Rule Them All?

For decades, the existence of the hydrophobic core, a region in the 3D structure of proteins where hydrophobic amino acids reside together, has been considered a general property in proteins. What we have found now may extend that model.

arXiv Machine Learning
Jun 16

Learning Topological Representations for Molecular Dynamics

arXiv:2606. 14737v1 Announce Type: cross Abstract: Molecular dynamics (MD) simulations generate trajectories in a high-dimensional configuration space whose analysis critically depends on molecular descriptors, typically handcrafted observables or learned kinetic embeddings.

By Dominik Geng, Florian Graf, Martin Uray, Roland Kwitt
arXiv Machine Learning
Aug 28

Interpreting Latent Protein Language Model Features with Geometric Annotations

The paper introduces a scalable method to interpret sparse autoencoder (SAE) features in the ESM-2 protein language model by leveraging geometrically inspired features of the protein α‑carbon backbone. Across 8M layers of ESM-2, a false discovery rate–controlled analysis shows that local geometry is significantly associated with many SAE features, revealing substructure within known biological labels and enabling annotation of unannotated metagenomic proteins. Ablation experiments demonstrate that removing these geometric features shifts ESM-2’s predicted contact maps toward the descriptor, linking mechanistic interpretability with structural biology.

By Siddharth Setlur, Djordje Mihajlovic, Darrick Lee
MIT News AI
Aug 27

Looking beyond natural sequences

A new machine‑learning framework is being developed to enhance the success rate of computational protein design. It deliberately moves away from reproducing sequences found in nature, aiming instead for novel designs that may perform better in practical applications.

By Lillian Eden | Department of Biology