arXiv Machine Learning

MSAlign: Aligning Molecule and Mass Spectra representations for Metabolite Identification

The paper introduces MSAlign, a lightweight model that aligns frozen foundation models for mass spectra (DreaMS) and molecules (MolDeBERTa) to improve metabolite identification from MS/MS spectra. It presents a unified framework for representation alignment and contrastive learning, demonstrates that a score fusion strategy further boosts performance at minimal cost, and addresses evaluation challenges by quantifying distribution shift in data splitting strategies. All resources, including datasets, splits, and code, are publicly released to promote reproducible research.

arXiv Machine Learning
Jul 23

Hypothesis-and-Refinement Learning of Organic Structures from Multimodal Spectroscopic Data

arXiv:2607. 19816v1 Announce Type: cross Abstract: Determining molecular structures from spectroscopic data remains fundamentally challenging because the inverse problem is intrinsically underdetermined: individual spectra are sparse, low-dimensional, and encode only partial structural evidence relative to the vast space of possible molecules.

By Chengchun Liu, Zhiyuan Yan, Li Yuan, Hao Li, Boxuan Zhao, Yonghong Tian, Bartosz A. Grzybowski, Fanyang Mo
arXiv Machine Learning
Jul 30

Data Fusion and Contrastive Alignment for Unconstrained IR Molecular Structure Elucidation

arXiv:2607. 26164v1 Announce Type: new Abstract: Automated molecular structure elucidation from infrared (IR) spectroscopy data has seen significant advancements in recent years, but its broad applicability is limited by a reliance on pre-determined chemical formulas provided as auxiliary model inputs.

By Ethan J. Mick, Campbell A. Sweet, Matthias J. Young, Derek T. Anderson
arXiv Machine Learning
Aug 20

Monroe: A Molecular Foundation Model for In-Context Probabilistic Inference

Monroe is a new molecular foundation model that improves upon existing models by pre‑training on over 81 million molecules from the PM6 quantum chemistry dataset, enhancing stereochemistry representation, and introducing novel training losses such as conformer denoising and embedding decorrelation. It also incorporates a prior‑data‑fitted model (TabPFN) for downstream in‑context prediction and demonstrates superior performance on Polaris benchmarks and activity cliff tests. Ablation studies show that the PFN‑based downstream approach can upgrade other models, producing state‑of‑the‑art variants MiniMol_PFN and CheMeleon_PFN.

By Blazej Banaszewski, Andrew W. Fitzgibbon
arXiv Machine Learning
Jun 19

MassSpecGym in the Wild: Uncovering and Correcting Evaluation Pitfalls in AI-Driven Molecule Discovery

arXiv:2606. 19624v1 Announce Type: new Abstract: Reliable benchmarking is critical for developing machine learning models for tandem mass spectrometry (MS/MS) based molecule discovery.

By Hongxuan Liu, Roman Bushuiev, Ivy Lightheart, Mrunali Manjrekar, Anton Bushuiev, Magdalena Lederbauer, Filip Jozefov, Yinkai Wang, Soha Hassoun, Josef Sivic, James Taylor, Runzhong Wang, David Healey, Tom\'a\v{s} Pluskal, Connor W. Coley
arXiv Machine Learning
Jul 28

FRIGID: Scaling Diffusion-Based Molecular Generation from Mass Spectra at Training and Inference Time

arXiv:2604. 16648v2 Announce Type: replace Abstract: Tandem mass spectrometry is prominent in scientific discovery workflows for identifying unknown small molecules, yet high-throughput structural elucidation remains challenging.

By Montgomery Bohde, Hongxuan Liu, Mrunali Manjrekar, Magdalena Lederbauer, Shuiwang Ji, Runzhong Wang, Connor W. Coley
arXiv AI
Aug 18

Multi-Agent Closed-Loop Reasoning for Organic Structure Elucidation from Multimodal Spectra

arXiv:2608. 14720v1 Announce Type: cross Abstract: Following the molecular discovery and synthesis revolutions, scalable automated structure elucidation from routine spectroscopic data remains an outstanding challenge.

By Bingsen Xue, Zhuojun Jiang, Jianhao Zhang, Mingcheng Gu, Yizhe Yuan, Yongtai Zhuo, Yifan Zhang, Li Wang, Ya Su, Yue Yuan, Jiang Liu, Xueqian Kong, Cheng Jin
arXiv Machine Learning
Sep 10

From Human Labels to Literature: Semi-Supervised Learning of NMR Chemical Shifts at Scale

The paper introduces a semi‑supervised framework that learns to predict nuclear magnetic resonance (NMR) chemical shifts from millions of literature‑extracted spectra without explicit atom‑level assignments. By treating the prediction as a permutation‑invariant set supervision problem, the authors show that optimal bipartite matching can be reduced to a sorting‑based loss, enabling stable large‑scale training. The resulting models outperform state‑of‑the‑art methods, generalize better to diverse molecules, and for the first time capture systematic solvent effects across common NMR solvents.

By Yongqi Jin, Yecheng Wang, Jun-jie Wang, Rong Zhu, Guolin Ke, Weinan E
arXiv AI
Aug 14

Simulation-to-real transfer learning for infrared spectroscopic chemical sensing and analysis from molecules to complex samples

arXiv:2608. 13341v1 Announce Type: cross Abstract: Infrared (IR) spectroscopy is widely used for chemical sensing, but extracting reliable chemical information from spectra remains challenging.

By Yusen Tan, Yixuan Chen, Zheng Fang, Pan Liu, Yifan Li, Qinyu Guo, Zhedong Lin, Yuqiang Li, Xiangxiang Zeng, Tong Wang, Jun Xia