arXiv:2607. 02140v1 Announce Type: new Abstract: Chemical language models (CLMs) are trained with linearized representations such as SMILES, yet it remains unclear which chemically meaningful substructures they encode.
By Anna Karnysheva, Dietrich Klakow, Ji-Ung Lee
The study evaluates four pretrained molecular language models on six virtual libraries covering drug discovery, organic materials, and catalysis. It finds that native embeddings vary widely in performance, while molecular fingerprints remain consistently strong. Fine‑tuning the models on library‑specific data markedly improves sample efficiency, with several adapted encoders outperforming others across all tasks.
By Henrik Wille, Luis-Finley Sch\"utz, Felix Strieth-Kalthoff
Mol-JEPA is a scalable multimodal framework that learns molecular world models by using modality masking instead of suboptimal perturbations. It incorporates diverse data such as molecular structures, cellular phenotypes, binding affinities, ADMET profiles, quantum chemistry simulations, and other drug‑discovery information. Benchmarks show that the representations it learns perform strongly, highlighting the benefit of embedding biochemical context via latent‑space prediction.
By Florian Rottach, Sebastian Schieferdecker, William Rudman, Randall Balestriero, Carsten Eickhoff
arXiv:2608. 02688v1 Announce Type: cross Abstract: Phenotypic drug discovery enables the discovery of functional relationships between molecular structures and cellular responses.
By Xuan Lin, Jingyu Sheng, Tengfei Ma, Li Sun, Dapeng Xiong
Large language models (LLMs) are widely applied across chemical tasks, such as molecular property prediction, which underpins drug discovery. Molecular LLMs represent a molecule through several modalities, notably a 1D SMILES sequence or a 2D molecular graph.
arXiv:2607. 01800v1 Announce Type: new Abstract: Large Language Models (LLMs) have recently shown promise in molecular discovery, yet a gap remains between their probabilistic nature over discrete sequential tokens and the rigid topological constraints of chemical space.
By Jiatong Li, Weida Wang, Changmeng Zheng, Shufei Zhang, Yatao Bian, Xiao-yong Wei, Qing Li
arXiv:2509. 22468v2 Announce Type: replace-cross Abstract: High-quality molecular representations are essential for property prediction and molecular design, yet large labeled datasets remain scarce.
By Boshra Ariguib, Mathias Niepert, Andrei Manolache
arXiv:2608. 10480v1 Announce Type: new Abstract: Large language models (LLMs) are widely applied across chemical tasks, such as molecular property prediction, which underpins drug discovery.
By Junwoo Park, Minyoung Shin, Cheol Soon Lee, Sujee Lee
arXiv:2607. 03007v1 Announce Type: cross Abstract: Recent advances in molecular large language models have led to strong performance on molecular understanding and generation tasks, yet these gains often come without reliable structural grounding.
By Wenda Wang, Jinjia Feng, Zhewei Wei
arXiv:2510.07289v2 Announce Type: replace
Abstract: Molecular graph representation learning is widely used in chemical and biomedical research. While pre-trained 2D graph encoders have demonstrated s...
By Xingtong Yu, Chang Zhou, Xinming Zhang, Yuan Fang
Local chemical perception and property reasoning are both essential for understanding how molecular structure determines properties. Current LLM-based chemical reasoning methods either receive SMILES/molecular images together with descriptions of local motifs, or reason directly from molecular images.
arXiv:2502. 07027v4 Announce Type: replace-cross Abstract: Molecular Relational Learning (MRL) is widely applied in natural sciences to predict relationships between molecular pairs by extracting structural features.
By Peiliang Zhang, Jingling Yuan, Qing Xie, Yongjun Zhu, Chao Che, Lin Li