arXiv Machine Learning By Parthasarathy Suryanarayanan, Susanta Das, Shreyans Sethi, Kenneth M. Merz, Jr., Joseph A. Morrone

Predicting Collision Cross Sections with GRACE: Geometric Residual Adduct Conditioning via Early-fusion

Read the original on arXiv Machine Learning →

The paper introduces GRACE, a 3D collision cross section (CCS) predictor that incorporates geometric residual adduct conditioning via early fusion. GRACE adapts a pretrained molecular geometry encoder with an adduct token and low‑rank attention adapters, achieving the lowest mean percentage differences on random, scaffold, and adduct‑sensitive splits of a curated dataset of over 9,000 experimental CCS records. Diagnostic analyses attribute its performance to residual learning that removes the dominant mass‑CCS trend and to early fusion that enhances adduct‑sensitive prediction.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jul 29

Beyond Predictive Accuracy: A Reliability-Aware Audit of Molecular Representations for Human Olfaction

arXiv:2607. 24848v1 Announce Type: cross Abstract: Pretrained molecular encoders are commonly evaluated through downstream prediction, but predictive accuracy alone does not establish that a learned representation captures reproducible scientific structure, adds information beyond strong conventional baselines, or transfers out of distribution.

By Kai Lun Huang (California State University, Fullerton), Wei Chieh Sun (University of Washington)
arXiv AI
Sep 17

Procedural Pretraining for Molecular Property Prediction

The paper proposes a three‑stage training pipeline that begins with procedural pretraining on abstract, procedurally generated data, followed by molecular pretraining on SMILES, and finally downstream fine‑tuning for molecular property prediction. Experiments show that procedural pretraining improves downstream performance—e.g., a 4.8% error reduction on Lipophilicity—especially when labeled data are scarce, and that the benefit peaks at an intermediate procedural training budget. Analysis indicates that transferable knowledge resides mainly in attention layers, while feed‑forward layers may over‑specialize.

By Moritz Friedemann, Zachary Shinnick, Philip Torr, Bruno Andreis
arXiv Machine Learning
Sep 22

MolSC: Leveraging Substituent Contributions to Enhance Fine-grained Molecular Understanding in LLMs

MolSC is a new dataset of 181,000 substituent-level examples that captures how attaching specific substituents to molecular scaffolds changes properties such as bioactivity and physicochemical descriptors. The authors also provide MolSC-Bench, a held‑out benchmark of 1,541 examples that are disjoint from MolSC at scaffold, substituent, and molecule levels. Experiments show that training molecular large language models on MolSC markedly improves their ability to predict substituent contributions, outperforming existing models on a range of downstream chemistry tasks.

By Hyuntae Park, Sooyeon Kim, Jiwon Park, SangKeun Lee