arXiv AI

Understanding Structural Representation in Foundation Models for Polymers

The paper introduces a chemical language foundation model that uses a SMILES‑based polymer graph representation (CPG) to encode polymer architecture and connectivity, addressing gaps in existing line notations. The model achieves strong performance across 30 polymer property benchmarks and demonstrates robustness to structural representation perturbations, with even chemically invalid SMILES sometimes matching state‑of‑the‑art results. Control experiments and attention analyses confirm that CPG offers meaningful advantages while highlighting the model’s ability to interpolate SMILES sequence space in a way that aligns loosely with chemical and architectural space.

arXiv AI
Sep 3

HiPoly: a hierarchical polymer-native AI framework for property prediction and generative design

HiPoly is a polymer-native AI framework that uses a three-level hierarchical graph architecture built on the G2RINS representation to process complete polymer descriptions. It encodes stochastic inter-monomer connectivity, composition, and molecular weight directly within its architecture, enabling end-to-end workflows from experimental data to property prediction, generative design, and physics-based validation. The framework achieves state-of-the-art accuracy for thermophysical properties of multi-component polymer systems and demonstrates generative design by discovering sustainable, PFAS-free alternatives with target surface-energy properties.

By Ge Sun, Gervasio Zaldivar, Yuan Tian, Gustavo Perez Lemus, Juhae Park, Dasha Safarian, Ming Han, Juan J. de Pablo
arXiv Machine Learning
1d ago

A Large Scale Investigation of Scaling Limits in Chemical Language Models

The paper reports a large-scale, compute-controlled study of Chemical Language Models (CLMs) involving over 30,000 experiments across different molecular representations, tokenizations, model sizes, datasets, and architectures. It finds clear scaling trends in pretraining loss but shows that these improvements do not translate into proportional gains in goal-directed molecular design, with chemical syntax saturating early while semantic properties develop more slowly. The authors release a new suite of models, NovoMolGen, that achieves state-of-the-art results in drug discovery tasks, highlighting a disconnect between representation learning and downstream design and calling for new pretraining paradigms that target chemical semantics.

By Roshan Balaji, Kamran Chitsaz, Quentin Fournier, Nirav Pravinbhai Bhatt, Sarath Chandar
arXiv Machine Learning
Jun 5

MolE-RAG: Molecular Structure-Enhanced Retrieval-Augmented Generation for Chemistry

arXiv:2606. 05693v1 Announce Type: new Abstract: Large language models (LLMs) have shown promise for molecular property prediction, but their ability to reason over chemical structures remains limited, as molecular representations such as SMILES differ substantially from the natural language on which LLMs are primarily trained.

By Joey Chan, Wonbin Kweon, Ashley Shin, Niharika Bhattacharjee, Pengcheng Jiang, Yue Guo, Jiawei Han
arXiv Machine Learning
Sep 1

Structural Hierarchy and Geometry in Molecular Representation Learning

The paper investigates how explicitly supervising molecular embeddings with a molecule’s Bemis‑Murcko scaffold influences representation learning. Experiments compare Euclidean and Lorentz contrastive objectives under two augmentation strengths, showing that scaffold‑supervised models consistently group molecules by identical and related scaffolds. These embeddings also enhance property prediction on several tasks, though the magnitude of improvement varies with the target property and the geometry used.

By David Sulu, Lorenzo Di Fruscia, Jana M. Weber