arXiv Machine Learning

Generative Molecular Morphing for Flexible-Size Design via Unbalanced Optimal Transport

arXiv:2606. 07239v1 Announce Type: new Abstract: The success of generative molecular design hinges on a model's steerability toward high-reward samples.

arXiv Machine Learning
Jul 23

Boltzmann-Expected Molecular Design with Decoupled Annealing Flows

arXiv:2607. 19519v1 Announce Type: cross Abstract: Most 3D properties relevant to molecular design, including free energies and shape descriptors, are $\textit{expectations}$ over the Boltzmann distribution over 3D configurations of a molecular graph.

By Selma Moqvist, Richard Beckmann, Ross Irwin, Roc\'io Mercado, Simon Olsson
arXiv Machine Learning
Jul 13

Autoregressive latent diffusion for 3D molecule generation

arXiv:2607. 09277v1 Announce Type: new Abstract: Three-dimensional (3D) molecule generation has been dominated by diffusion models, which achieve strong generation quality but typically require the molecular size to be specified a priori.

By Federico Ottomano, Gaopeng Ren, Yingzhen Li, Kim E. Jelfs, Alex M. Ganose
arXiv AI
Sep 7

NEAT-POCKET: Pocket-Conditioned Autoregressive 3D Molecular Generation with a Neighborhood-Guided Set Transformer

NEAT-POCKET is a pocket‑conditioned extension of the autoregressive NEAT model that generates 3D molecules atom by atom within protein binding pockets, maintaining atom permutation invariance and explicitly modeling hydrogen atoms. It outperforms existing baselines on the CrossDocked and SPINDR datasets, achieving competitive structure‑based generation performance while sampling significantly faster. The model also supports pocket‑conditioned fragment completion, a capability directly useful for lead optimization and scaffold elaboration in drug design.

By Roxane Axel Jacob, Daniel Rose, Thierry Langer, Johannes Kirchmair
Hugging Face Trending Papers
Sep 8

Fixed-Dimensional Latent Flow for Generating Variable-Size 3D Molecules

The paper introduces Equivariant-Free Transformer-Autoencoded Latent Flow Matching (EF‑TALFM), a two‑stage generative framework that uses a single fixed‑dimensional latent vector to produce variable‑size 3D molecules. The first stage samples the latent vector via flow matching, and the second stage employs an autoregressive Transformer decoder that determines molecule size while generating atom types, coordinates, and chemical states. EF‑TALFM outperforms prior methods on the PCQM4Mv2 benchmark, achieving higher uniqueness, novelty, and computational throughput, and its internal ranking improves the hit rate for target HOMO–LUMO gaps while maintaining novelty.

arXiv AI
Jul 21

Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial Constraints

arXiv:2607. 18144v1 Announce Type: cross Abstract: Structure-based drug design (SBDD) leverages the 3D structure of protein targets, often complemented by other spatial constraints, to generate candidate binding molecules.

By Thomas MacDougall, Maksim Kuznetsov, Roman Schutski, Rim Shayakhmetov, Maxim Malkov, Vladimir Aladinskiy, Alex Aliper, Alex Zhavoronkov
arXiv Machine Learning
Sep 23

Transport-Coupled Bayesian Flows for Molecular Graph Generation

Transport-Coupled Bayesian Flows for Molecular Graph Generation (TopBF) addresses a key mismatch in existing diffusion models for molecular graph generation by eliminating the need for hard discretization during sampling. The framework generates graphs directly in continuous parameter distributions, learns graph topology via a Quasi-Wasserstein optimal‑transport coupling with geodesic costs, and enables property‑conditioned generation without retraining. Experiments on QM9 and ZINC250k show that TopBF achieves higher structural fidelity and more efficient generation compared to prior methods.

By Yida Xiong, Jiameng Chen, Kun Li, Hongzhi Zhang, Xiantao Cai, Lei Lei, Wenbin Hu
arXiv AI
2d ago

CODesign: Consistency from Data to Trajectory in All-Atom Protein Binder Co-Design

CODesign is a co-design framework that jointly generates protein sequences and structures to improve consistency between them. It introduces a large consistency‑distilled dataset of about 105,000 dimers and employs a multimodal joint flow model with a consistency‑aware resampling strategy to iteratively refine sequences and side chains. The approach achieves state‑of‑the‑art in silico success rates for protein‑ and ligand‑target binder design, with ablation studies showing a 70.9% performance boost from the distilled dataset and further gains from the resampling mechanism.

By Yuanle Mo, Bo Qiang, Haitao Lin, Qinghan Wang, Gang Du, Odin Zhang, Pheng Ann Heng