Scalable Peptide Design via Memory-Efficient Equivariant Transformer
arXiv:2606. 25006v1 Announce Type: new Abstract: Target-specific peptide design requires sequence and structure co-design under full atom geometric constraints.
Target-specific peptide design requires sequence and structure co-design under full atom geometric constraints. Latent generative frameworks offer an effective route for this problem by compressing fine grained atomic structures into block level latent representations and performing conditional generation in a compact latent space.
arXiv:2606. 25006v1 Announce Type: new Abstract: Target-specific peptide design requires sequence and structure co-design under full atom geometric constraints.
arXiv:2607. 03787v1 Announce Type: new Abstract: Accurately modeling biomolecular interactions is a central bottleneck in biology and therapeutic discovery.
arXiv:2607. 09998v1 Announce Type: new Abstract: Macrocyclic peptides are an increasingly important therapeutic modality, but existing computational methods for modeling their structures and properties are limited in scope and do not generalize well across the synthetically accessible chemical space.
CODesign is a co-design framework that jointly generates protein sequences and structures to improve consistency between them. It introduces a large consistency‑distilled dataset of about 105,000 dimers and employs a multimodal joint flow model with a consistency‑aware resampling strategy to iteratively refine sequences and side chains. The approach achieves state‑of‑the‑art in silico success rates for protein‑ and ligand‑target binder design, with ablation studies showing a 70.9% performance boost from the distilled dataset and further gains from the resampling mechanism.
The paper introduces Equivariant-Free Transformer-Autoencoded Latent Flow Matching (EF‑TALFM), a two‑stage generative framework that uses a single fixed‑dimensional latent vector to produce variable‑size 3D molecules. The first stage samples the latent vector via flow matching, and the second stage employs an autoregressive Transformer decoder that determines molecule size while generating atom types, coordinates, and chemical states. EF‑TALFM outperforms prior methods on the PCQM4Mv2 benchmark, achieving higher uniqueness, novelty, and computational throughput, and its internal ranking improves the hit rate for target HOMO–LUMO gaps while maintaining novelty.
arXiv:2609.08333v1 Announce Type: cross Abstract: In molecular discovery, molecule size is coupled to composition, structure, and other target properties. Yet most 3D generators require molecule size...
arXiv:2603. 19636v2 Announce Type: replace Abstract: Accurate RNA structure modeling remains difficult because RNA backbones are highly flexible, non-canonical interactions are prevalent, and experimentally determined 3D structures are comparatively scarce.
arXiv:2607. 12380v1 Announce Type: new Abstract: Small molecules, crystals, and proteins all reduce to atoms in 3D space, yet their generative pipelines remain fragmented across domains, each with its Small molecules, crystals, and proteins all reduce to atoms in 3D space, yet their generative pipelines remain fragmented across domains, each with its own graph, equivariant, or frame-based architecture.
TorchCraft is a unified binder‑design framework that optimizes sequence logits using a frozen all‑atom structure predictor. It integrates confidence, contact, geometric, and sequence‑prior objectives within TorchFold to design minibinders, framework‑conditioned VHHs, cyclic peptides, and ligand‑binding proteins. Using pretrained AlphaFold 3 weights, TorchCraft produced experimentally validated binders across four targets without post‑hoc redesign, and computational tests confirmed its applicability to cyclic peptides and ligand‑conditioned pocket design.
The paper reports a large-scale, compute-controlled study of Chemical Language Models (CLMs) involving over 30,000 experiments across different molecular representations, tokenizations, model sizes, datasets, and architectures. It finds clear scaling trends in pretraining loss but shows that these improvements do not translate into proportional gains in goal-directed molecular design, with chemical syntax saturating early while semantic properties develop more slowly. The authors release a new suite of models, NovoMolGen, that achieves state-of-the-art results in drug discovery tasks, highlighting a disconnect between representation learning and downstream design and calling for new pretraining paradigms that target chemical semantics.
NEAT-POCKET is a pocket‑conditioned extension of the autoregressive NEAT model that generates 3D molecules atom by atom within protein binding pockets, maintaining atom permutation invariance and explicitly modeling hydrogen atoms. It outperforms existing baselines on the CrossDocked and SPINDR datasets, achieving competitive structure‑based generation performance while sampling significantly faster. The model also supports pocket‑conditioned fragment completion, a capability directly useful for lead optimization and scaffold elaboration in drug design.
arXiv:2607. 01105v1 Announce Type: new Abstract: We present SynLaD, a latent diffusion framework for small-molecule generation that unifies ligand-based drug design objectives (what to make) with synthetic accessibility (how to make it).