arXiv Machine Learning

Probing and steering biology across Boltz-1s trunk-diffusion boundary

arXiv:2608. 11475v1 Announce Type: cross Abstract: AlphaFold3-class structure predictors pair a representational trunk, which processes sequence and context, with a diffusion module, which generates atomic coordinates.

arXiv Machine Learning
Jun 30

Inference-time optimization for experiment-grounded protein ensemble generation

arXiv:2602. 24007v3 Announce Type: replace-cross Abstract: Protein function relies on dynamic conformational ensembles, yet current generative models like AlphaFold3 often fail to produce ensembles that match experimental data.

By Advaith Maddipatla, Anar Rzayev, Marco Pegoraro, Martin Pacesa, Paul Schanda, Ailie Marx, Sanketh Vedula, Alex M. Bronstein
arXiv Machine Learning
Jun 9

Few-step Cofolding with All-Atom Flow Maps

arXiv:2606. 08375v1 Announce Type: new Abstract: All-atom generative modeling of 3D biomolecular complexes has emerged as the dominant paradigm for predicting the structure of proteins and protein-ligand systems.

By Gianluca Scarpellini, Ron Shprints, Peter Holderrieth, Juno Nam, Pranav Murugan, Rafael G\'omez-Bombarelli, Tommi Jaakola, Maruan Al-Shedivat, Nicholas Matthew Boffi, Avishek Joey Bose
arXiv Machine Learning
Jun 29

PairSAE: Mechanistic Interpretability from Pair Representations in Protein Co-Folding

arXiv:2606. 27440v1 Announce Type: new Abstract: Foundation models for structural biology have achieved remarkable performance in predicting biomolecular structure and show promise for the design of proteins and small molecules.

By Giosue Migliorini, Aristofanis Rontogiannis, Grigori Guitchounts, Nicholas Franklin, Axel Elaldi, Olivia Viessmann
arXiv Machine Learning
1d ago

A Large Scale Investigation of Scaling Limits in Chemical Language Models

The paper reports a large-scale, compute-controlled study of Chemical Language Models (CLMs) involving over 30,000 experiments across different molecular representations, tokenizations, model sizes, datasets, and architectures. It finds clear scaling trends in pretraining loss but shows that these improvements do not translate into proportional gains in goal-directed molecular design, with chemical syntax saturating early while semantic properties develop more slowly. The authors release a new suite of models, NovoMolGen, that achieves state-of-the-art results in drug discovery tasks, highlighting a disconnect between representation learning and downstream design and calling for new pretraining paradigms that target chemical semantics.

By Roshan Balaji, Kamran Chitsaz, Quentin Fournier, Nirav Pravinbhai Bhatt, Sarath Chandar
arXiv Machine Learning
Aug 28

Interpreting Latent Protein Language Model Features with Geometric Annotations

The paper introduces a scalable method to interpret sparse autoencoder (SAE) features in the ESM-2 protein language model by leveraging geometrically inspired features of the protein α‑carbon backbone. Across 8M layers of ESM-2, a false discovery rate–controlled analysis shows that local geometry is significantly associated with many SAE features, revealing substructure within known biological labels and enabling annotation of unannotated metagenomic proteins. Ablation experiments demonstrate that removing these geometric features shifts ESM-2’s predicted contact maps toward the descriptor, linking mechanistic interpretability with structural biology.

By Siddharth Setlur, Djordje Mihajlovic, Darrick Lee
arXiv Machine Learning
Jul 15

SinAE: A Single-Architecture Flow-Matching Autoencoder for Cross-Domain Atomic Systems

arXiv:2607. 12380v1 Announce Type: new Abstract: Small molecules, crystals, and proteins all reduce to atoms in 3D space, yet their generative pipelines remain fragmented across domains, each with its Small molecules, crystals, and proteins all reduce to atoms in 3D space, yet their generative pipelines remain fragmented across domains, each with its own graph, equivariant, or frame-based architecture.

By Yuxuan Ren, Fan Yang, Jianhua Yao, Yatao Bian