arXiv:2606. 03057v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for molecular tasks, but it remains unclear which molecular representation to use.
By Arun Raja, Garrett M. Morris, Kian Ming A. Chai
arXiv:2606. 05693v1 Announce Type: new Abstract: Large language models (LLMs) have shown promise for molecular property prediction, but their ability to reason over chemical structures remains limited, as molecular representations such as SMILES differ substantially from the natural language on which LLMs are primarily trained.
By Joey Chan, Wonbin Kweon, Ashley Shin, Niharika Bhattacharjee, Pengcheng Jiang, Yue Guo, Jiawei Han
MolEmb is a lightweight framework that adapts multimodal large language models (MLLMs) to serve as general molecular embedding models. By aligning molecular profiles with textual descriptions in a shared embedding space using a bidirectional contrastive objective, MolEmb produces embeddings conditioned on both a molecular profile and a natural‑language semantic context. The model performs competitively on molecular property prediction and enables cross‑modal molecule‑text retrieval, while the newly introduced MolCAR benchmark demonstrates that context‑aware molecular embedding is largely a data property of the supervision.
By Xinjian Zhao, Xiangru Jian, Yaoyao Xu, Xiaozhuang Song, Wei Pang, Lei Bai, Tianshu Yu
arXiv:2603. 25062v2 Announce Type: replace Abstract: Autoregressive molecular models assign probability to molecular serializations even though chemical identity is invariant to serialization.
By Xinyu Wang, Fei Dou, Jinbo Bi, Minghu Song
The paper reports a large-scale, compute-controlled study of Chemical Language Models (CLMs) involving over 30,000 experiments across different molecular representations, tokenizations, model sizes, datasets, and architectures. It finds clear scaling trends in pretraining loss but shows that these improvements do not translate into proportional gains in goal-directed molecular design, with chemical syntax saturating early while semantic properties develop more slowly. The authors release a new suite of models, NovoMolGen, that achieves state-of-the-art results in drug discovery tasks, highlighting a disconnect between representation learning and downstream design and calling for new pretraining paradigms that target chemical semantics.
By Roshan Balaji, Kamran Chitsaz, Quentin Fournier, Nirav Pravinbhai Bhatt, Sarath Chandar
The paper introduces Align-React, a chemical reaction representation learning framework that incorporates atomic correspondence between reactants and products, an adapter for embedding reaction conditions, and a Reaction-Center-Aware attention mechanism. These components enable the model to capture precise molecular transformations and focus on critical functional groups, leading to improved performance across a variety of organic reaction tasks. The framework outperforms existing architectures on most benchmark datasets.
By Kaipeng Zeng, Xianbin Liu, Yu Zhang, Xiaokang Yang, Yaohui Jin, Yanyan Xu
arXiv:2602. 02320v4 Announce Type: replace-cross Abstract: Molecular function is largely determined by structure.
By Feiyang Cai, Guijuan He, Yi Hu, Jingjing Wang, Joshua Luo, Tianyu Zhu, Srikanth Pilla, Gang Li, Ling Liu, Feng Luo
ChemMLLM is a unified chemical multimodal large language model designed for molecule understanding and generation across text, SMILES strings, and images. The authors curated five multimodal tasks and benchmarked ChemMLLM against leading general MLLMs, chemical LLMs, and specialized models, finding it outperforms general-purpose MLLMs and matches specialized models on all tasks. The study demonstrates that a single foundation model can handle diverse cross‑modal chemical tasks, including image generation, enabling more intuitive visual human‑AI interaction.
By Qian Tan, Di Zhang, Ben Gao, Peng Xia, Wanhao Liu, Shufei Zhang, Wanli Ouyang, Lei Bai, Yuqiang Li, Tianfan Fu
arXiv:2606. 11382v1 Announce Type: new Abstract: Deep learning models facilitate the discovery of molecules with tailored properties among billions of candidate compounds.
By Emily Nguyen, Yongchan Hong, Harsh Toshniwal, Yan Liu, Andreas Luttens
arXiv:2606. 11508v1 Announce Type: new Abstract: Accurate prediction of absorption, distribution, metabolism, and excretion (ADME) properties is critical to drug discovery, but remains challenging because ADME endpoints are noisy, interdependent, and often data-limited.
By Yifan Xue, Srimukh Prasad Veccham, Saee Paliwal, Tyler Shimko, Micha Livne
The study investigates how document segmentation and chunk representation affect retrieval-augmented generation (RAG) for chemistry texts. Using the ChemQuests corpus, the authors benchmark 41 embedding models and evaluate them across five chunking strategies, seven chunk sizes, and various overlap settings. They find that embedding choice has the largest impact, with models like E5, BGE, and Nomic performing best, and recommend medium-to-large chunks with fixed-token, recursive-token, or hierarchical-section chunking and low overlap for effective chemistry-aware RAG.
By Mahmoud Amiri, Thomas Bocklitz
arXiv:2605. 16823v2 Announce Type: replace Abstract: Large language models succeed by combining large-scale pretraining with meaningful discrete tokens.
By Takayuki Kimura