arXiv:2607. 19406v1 Announce Type: new Abstract: Structural elucidation from Nuclear Magnetic Resonance (NMR) data remains a fundamental bottleneck across chemistry, materials science, and biology.
By Irina Espejo Morales, Damon Hinz, Marvin Alberts, Geraud Krawezik, Haewon Jeong, Shirley Ho
arXiv:2512. 18531v2 Announce Type: replace-cross Abstract: One-dimensional NMR spectroscopy is one of the most widely used techniques for the characterization of organic compounds and natural products.
By Frank Hu, Jonathan M. Tubb, Dimitris Argyropoulos, Sergey Golotvin, Mikhail Elyashberg, Grant M. Rotskoff, Matthew W. Kanan, Thomas E. Markland
arXiv:2603. 23101v3 Announce Type: replace Abstract: Intelligent spectroscopy serves as a pivotal element in AI-driven closed-loop scientific discovery, functioning as the critical bridge between matter structure and artificial intelligence.
By Yutang Ge, Yaning Cui, Hanzheng Li, Jun-Jie Wang, Fanjie Xu, Jinhan Dong, Yongqi Jin, Dongxu Cui, Peng Jin, Guojiang Zhao, Hengxing Cai, Tianci Yangfeng, Xueqing Chen, Hongshuai Wang, Rong Zhu, Linfeng Zhang, Xiaohong Ji, Zhifeng Gao
arXiv:2606. 30961v1 Announce Type: cross Abstract: Advances in deep learning architectures and representations have enabled ML-driven chemical property prediction, but state-of-the-art (SOTA) models have remained largely confined to independent codebases and lack support for diverse chemical species.
By Jacob W. Toney, Samir Darouich, Yiran Wang, Aaron G. Garrison, Johannes K\"astner, Heather J. Kulik
arXiv:2512. 19733v3 Announce Type: replace-cross Abstract: Molecular structure elucidation from spectroscopic data is a long-standing challenge in Chemistry, traditionally requiring expert interpretation.
By Federico Ottomano, Yingzhen Li, Alex M. Ganose
The paper introduces MSAlign, a lightweight model that aligns frozen foundation models for mass spectra (DreaMS) and molecules (MolDeBERTa) to improve metabolite identification from MS/MS spectra. It presents a unified framework for representation alignment and contrastive learning, demonstrates that a score fusion strategy further boosts performance at minimal cost, and addresses evaluation challenges by quantifying distribution shift in data splitting strategies. All resources, including datasets, splits, and code, are publicly released to promote reproducible research.
By Paul Krzakala, Gabriel Melo, Camille Lan\c{c}on, Charlotte Laclau, R\'emi Flamary, Etienne Th\'evenot, Florence d'Alch\'e-Buc
Monroe is a new molecular foundation model that improves upon existing models by pre‑training on over 81 million molecules from the PM6 quantum chemistry dataset, enhancing stereochemistry representation, and introducing novel training losses such as conformer denoising and embedding decorrelation. It also incorporates a prior‑data‑fitted model (TabPFN) for downstream in‑context prediction and demonstrates superior performance on Polaris benchmarks and activity cliff tests. Ablation studies show that the PFN‑based downstream approach can upgrade other models, producing state‑of‑the‑art variants MiniMol_PFN and CheMeleon_PFN.
By Blazej Banaszewski, Andrew W. Fitzgibbon
arXiv:2607. 19816v1 Announce Type: cross Abstract: Determining molecular structures from spectroscopic data remains fundamentally challenging because the inverse problem is intrinsically underdetermined: individual spectra are sparse, low-dimensional, and encode only partial structural evidence relative to the vast space of possible molecules.
By Chengchun Liu, Zhiyuan Yan, Li Yuan, Hao Li, Boxuan Zhao, Yonghong Tian, Bartosz A. Grzybowski, Fanyang Mo
arXiv:2606. 11508v1 Announce Type: new Abstract: Accurate prediction of absorption, distribution, metabolism, and excretion (ADME) properties is critical to drug discovery, but remains challenging because ADME endpoints are noisy, interdependent, and often data-limited.
By Yifan Xue, Srimukh Prasad Veccham, Saee Paliwal, Tyler Shimko, Micha Livne
arXiv:2608. 13341v1 Announce Type: cross Abstract: Infrared (IR) spectroscopy is widely used for chemical sensing, but extracting reliable chemical information from spectra remains challenging.
By Yusen Tan, Yixuan Chen, Zheng Fang, Pan Liu, Yifan Li, Qinyu Guo, Zhedong Lin, Yuqiang Li, Xiangxiang Zeng, Tong Wang, Jun Xia
arXiv:2608. 07610v1 Announce Type: cross Abstract: Overlapping peaks and sample-dependent chemical shift variability prevent reliable metabolite recovery from complex biological spectra.
By Jesper L{\o}ve Hinrich, Pia Susan Mayer, Bekzod Khakimov, S{\o}ren Balling Engelsen, Morten M{\o}rup
arXiv:2608. 03855v1 Announce Type: new Abstract: Transformer models have revolutionized natural language processing (NLP), and text-based molecular representations like SMILES have successfully extended these architectures to chemistry.
By David Ming Segura, Jeremy Goumaz, Joshua W. Sin, Bojana Rankovi\'c, Philippe Schwaller