arXiv AI

Higher-Order Molecular Grammars for Generative and Foundation Models in Chemistry

arXiv Machine Learning
5d ago

WEECFP-SuRGE: A Position-Aware Substructure Encoding Method for Molecular Property Prediction

WEECFP-SuRGE introduces a position‑aware substructure encoding method that combines tokenized hierarchical Morgan fingerprints with graph‑distance‑dependent rotations applied at the input and within transformer self‑attention. The approach captures local chemistry, long‑range interactions, and molecular topology without requiring external pretraining or 3‑D conformer generation. Benchmarks on MoleculeNet and the Therapeutic Data Commons ADMET datasets show competitive performance, and a reconstruction procedure correctly identifies constitutional isomers for 92.6% of a 4,200‑molecule library.

By Robert Epps
arXiv Machine Learning
Jul 8

Multimodal Molecular Representation Learning with Graph Neural Networks, Deep & Cross Networks, and SMILES Embeddings

arXiv:2607. 05736v1 Announce Type: new Abstract: Molecular property prediction often relies on isolated data modalities, where continuous 3D graph neural networks (GNNs) struggle to efficiently capture long-range topological dependencies and exact macroscopic heuristics.

By Qiwei Han, Chi Zhou, Ruobing Wang, Zheng Ma
arXiv Machine Learning
1d ago

A Large Scale Investigation of Scaling Limits in Chemical Language Models

The paper reports a large-scale, compute-controlled study of Chemical Language Models (CLMs) involving over 30,000 experiments across different molecular representations, tokenizations, model sizes, datasets, and architectures. It finds clear scaling trends in pretraining loss but shows that these improvements do not translate into proportional gains in goal-directed molecular design, with chemical syntax saturating early while semantic properties develop more slowly. The authors release a new suite of models, NovoMolGen, that achieves state-of-the-art results in drug discovery tasks, highlighting a disconnect between representation learning and downstream design and calling for new pretraining paradigms that target chemical semantics.

By Roshan Balaji, Kamran Chitsaz, Quentin Fournier, Nirav Pravinbhai Bhatt, Sarath Chandar
arXiv Machine Learning
Aug 10

How Molecular Generative Models Organize Molecular Identity

arXiv:2608. 06956v1 Announce Type: new Abstract: Generative models for matter are often evaluated as samplers over output representations, and their latent spaces are commonly used as proxies for navigating chemical space.

By Raul Ortega-Ochoa, Tejs Vegge, Jens S. Bakander, Luis Mantilla Calderon, Alan Aspuru-Guzik, Tonio Buonassisi
arXiv AI
2d ago

CODesign: Consistency from Data to Trajectory in All-Atom Protein Binder Co-Design

CODesign is a co-design framework that jointly generates protein sequences and structures to improve consistency between them. It introduces a large consistency‑distilled dataset of about 105,000 dimers and employs a multimodal joint flow model with a consistency‑aware resampling strategy to iteratively refine sequences and side chains. The approach achieves state‑of‑the‑art in silico success rates for protein‑ and ligand‑target binder design, with ablation studies showing a 70.9% performance boost from the distilled dataset and further gains from the resampling mechanism.

By Yuanle Mo, Bo Qiang, Haitao Lin, Qinghan Wang, Gang Du, Odin Zhang, Pheng Ann Heng
arXiv AI
Aug 19

Domain-Adapted Molecular Language Models for Efficient Search of Make-on-Demand Libraries

The study evaluates four pretrained molecular language models on six virtual libraries covering drug discovery, organic materials, and catalysis. It finds that native embeddings vary widely in performance, while molecular fingerprints remain consistently strong. Fine‑tuning the models on library‑specific data markedly improves sample efficiency, with several adapted encoders outperforming others across all tasks.

By Henrik Wille, Luis-Finley Sch\"utz, Felix Strieth-Kalthoff
arXiv Machine Learning
Jul 14

Vilya-1: An all-atom foundation model for macrocycle structure prediction and design

arXiv:2607. 09998v1 Announce Type: new Abstract: Macrocyclic peptides are an increasingly important therapeutic modality, but existing computational methods for modeling their structures and properties are limited in scope and do not generalize well across the synthetically accessible chemical space.

By Vilya Research, :, Pascal Sturmfels, Milad Salem, Naozumi Hiranuma, Stephen Rettie, Xiaoliang Pan, Benjamin D. Sellers, Adam P. Moyer, Patrick J. Salveson, Ivan Anishchanka
arXiv Machine Learning
Aug 28

Packora: Systematic Design for Generative Molecular Crystal Structure Prediction

Packora is a flow-based generative model designed for molecular crystal structure prediction (CSP). It jointly predicts atomic coordinates and lattice parameters from molecular graphs, supporting multi-component and organometallic crystals and allowing conditioning on conformers, stereochemistry, and space-group data. In evaluations inspired by the CCDC CSP blind test, Packora outperforms baselines on generation and ranking benchmarks, achieving superior matched-budget coverage, higher experimental-form recovery, lower ranks, and faster convergence.

By Nayoung Kim, Kiyoung Seong, Sungsoo Ahn