arXiv Machine Learning

Scalable Peptide Design via Memory-Efficient Equivariant Transformer

arXiv:2606. 25006v1 Announce Type: new Abstract: Target-specific peptide design requires sequence and structure co-design under full atom geometric constraints.

arXiv Machine Learning
Jul 14

Vilya-1: An all-atom foundation model for macrocycle structure prediction and design

arXiv:2607. 09998v1 Announce Type: new Abstract: Macrocyclic peptides are an increasingly important therapeutic modality, but existing computational methods for modeling their structures and properties are limited in scope and do not generalize well across the synthetically accessible chemical space.

By Vilya Research, :, Pascal Sturmfels, Milad Salem, Naozumi Hiranuma, Stephen Rettie, Xiaoliang Pan, Benjamin D. Sellers, Adam P. Moyer, Patrick J. Salveson, Ivan Anishchanka
Hugging Face Trending Papers
Sep 8

Fixed-Dimensional Latent Flow for Generating Variable-Size 3D Molecules

The paper introduces Equivariant-Free Transformer-Autoencoded Latent Flow Matching (EF‑TALFM), a two‑stage generative framework that uses a single fixed‑dimensional latent vector to produce variable‑size 3D molecules. The first stage samples the latent vector via flow matching, and the second stage employs an autoregressive Transformer decoder that determines molecule size while generating atom types, coordinates, and chemical states. EF‑TALFM outperforms prior methods on the PCQM4Mv2 benchmark, achieving higher uniqueness, novelty, and computational throughput, and its internal ranking improves the hit rate for target HOMO–LUMO gaps while maintaining novelty.

arXiv AI
2d ago

CODesign: Consistency from Data to Trajectory in All-Atom Protein Binder Co-Design

CODesign is a co-design framework that jointly generates protein sequences and structures to improve consistency between them. It introduces a large consistency‑distilled dataset of about 105,000 dimers and employs a multimodal joint flow model with a consistency‑aware resampling strategy to iteratively refine sequences and side chains. The approach achieves state‑of‑the‑art in silico success rates for protein‑ and ligand‑target binder design, with ablation studies showing a 70.9% performance boost from the distilled dataset and further gains from the resampling mechanism.

By Yuanle Mo, Bo Qiang, Haitao Lin, Qinghan Wang, Gang Du, Odin Zhang, Pheng Ann Heng
arXiv Machine Learning
1d ago

A Large Scale Investigation of Scaling Limits in Chemical Language Models

The paper reports a large-scale, compute-controlled study of Chemical Language Models (CLMs) involving over 30,000 experiments across different molecular representations, tokenizations, model sizes, datasets, and architectures. It finds clear scaling trends in pretraining loss but shows that these improvements do not translate into proportional gains in goal-directed molecular design, with chemical syntax saturating early while semantic properties develop more slowly. The authors release a new suite of models, NovoMolGen, that achieves state-of-the-art results in drug discovery tasks, highlighting a disconnect between representation learning and downstream design and calling for new pretraining paradigms that target chemical semantics.

By Roshan Balaji, Kamran Chitsaz, Quentin Fournier, Nirav Pravinbhai Bhatt, Sarath Chandar
arXiv AI
Sep 18

TorchCraft: Unified binder design by inverting an all-atom structure predictor

TorchCraft is a unified binder‑design framework that optimizes sequence logits using a frozen all‑atom structure predictor. It integrates confidence, contact, geometric, and sequence‑prior objectives within TorchFold to design minibinders, framework‑conditioned VHHs, cyclic peptides, and ligand‑binding proteins. Using pretrained AlphaFold 3 weights, TorchCraft produced experimentally validated binders across four targets without post‑hoc redesign, and computational tests confirmed its applicability to cyclic peptides and ligand‑conditioned pocket design.

By TorchCraft Team, Yu Liu, Zhouhanyu Shen, Zhengyi Li, Xikun Huang, Jiaqi Liu, Shuxian Gao, Qilin Yu, Xiayan Qin, Yucheng Zhang, Mingchen Chen
arXiv Machine Learning
Jul 15

SinAE: A Single-Architecture Flow-Matching Autoencoder for Cross-Domain Atomic Systems

arXiv:2607. 12380v1 Announce Type: new Abstract: Small molecules, crystals, and proteins all reduce to atoms in 3D space, yet their generative pipelines remain fragmented across domains, each with its Small molecules, crystals, and proteins all reduce to atoms in 3D space, yet their generative pipelines remain fragmented across domains, each with its own graph, equivariant, or frame-based architecture.

By Yuxuan Ren, Fan Yang, Jianhua Yao, Yatao Bian
arXiv AI
Sep 7

NEAT-POCKET: Pocket-Conditioned Autoregressive 3D Molecular Generation with a Neighborhood-Guided Set Transformer

NEAT-POCKET is a pocket‑conditioned extension of the autoregressive NEAT model that generates 3D molecules atom by atom within protein binding pockets, maintaining atom permutation invariance and explicitly modeling hydrogen atoms. It outperforms existing baselines on the CrossDocked and SPINDR datasets, achieving competitive structure‑based generation performance while sampling significantly faster. The model also supports pocket‑conditioned fragment completion, a capability directly useful for lead optimization and scaffold elaboration in drug design.

By Roxane Axel Jacob, Daniel Rose, Thierry Langer, Johannes Kirchmair