The study investigates how document segmentation and chunk representation affect retrieval-augmented generation (RAG) for chemistry texts. Using the ChemQuests corpus, the authors benchmark 41 embedding models and evaluate them across five chunking strategies, seven chunk sizes, and various overlap settings. They find that embedding choice has the largest impact, with models like E5, BGE, and Nomic performing best, and recommend medium-to-large chunks with fixed-token, recursive-token, or hierarchical-section chunking and low overlap for effective chemistry-aware RAG.
By Mahmoud Amiri, Thomas Bocklitz
arXiv:2608. 03855v1 Announce Type: new Abstract: Transformer models have revolutionized natural language processing (NLP), and text-based molecular representations like SMILES have successfully extended these architectures to chemistry.
By David Ming Segura, Jeremy Goumaz, Joshua W. Sin, Bojana Rankovi\'c, Philippe Schwaller
arXiv:2602. 02320v4 Announce Type: replace-cross Abstract: Molecular function is largely determined by structure.
By Feiyang Cai, Guijuan He, Yi Hu, Jingjing Wang, Joshua Luo, Tianyu Zhu, Srikanth Pilla, Gang Li, Ling Liu, Feng Luo
arXiv:2606. 05693v1 Announce Type: new Abstract: Large language models (LLMs) have shown promise for molecular property prediction, but their ability to reason over chemical structures remains limited, as molecular representations such as SMILES differ substantially from the natural language on which LLMs are primarily trained.
By Joey Chan, Wonbin Kweon, Ashley Shin, Niharika Bhattacharjee, Pengcheng Jiang, Yue Guo, Jiawei Han
arXiv:2608. 03525v3 Announce Type: replace Abstract: In organic chemistry papers and patents, molecular structures, reaction schemes, and experimental conditions are often presented as molecular structure depictions, reaction diagrams, and complex tables or figures.
By Haote Yang, Jiang Wu, Jingchao Wang, Xingjian Wei, Lixin Ma, Linye Li, Chen Zhu, Xiaolong Wu, Yuheng Lu, Ziran Zhu, Junyuan Gao, Lingli Ge, Yuan Xu, Huijie Ao, QianQian Wu, Dechen Lin, Huaiyu Gu, Lu Chen, Shengxin Lu, ShaSha Wang, Yuanyuan Cao, Zhejia Yu, Ruijie Zhang, Zimai Tian, Jiaxing Sun, Yinfan Wang, Jiahe Song, Chuang Wang, Yubin Wang, Rui Nie, Hao Zheng, Bowen Jiang, Hongbin Lai, Yifan He, Chengjin Liu, Tingting Zhang, Liqun Wei, Lijun Wu, Bin Wang, Yuqiang Li, Guangyu Wang, Wei Li, Bowen Zhou, Dahua Lin, Conghui He
ChemVTS-Bench is a domain-authentic benchmark that evaluates Visual‑Textual‑Symbolic reasoning in multimodal large language models for chemistry. It presents diverse chemical problems—organic molecules, inorganic materials, and 3D crystal structures—in three input modes: visual-only, visual‑text hybrid, and SMILES-based symbolic. The benchmark includes an automated agent workflow for inference, answer verification, and failure diagnosis, and shows that visual-only inputs and structural chemistry remain challenging for current models.
By Zhiyuan Huang, Baichuan Yang, Zikun He, Yanhong Wu, Fang Hongyu, Zhenhe Liu, Lin Dongsheng, Bing Su
arXiv:2606. 03057v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for molecular tasks, but it remains unclear which molecular representation to use.
By Arun Raja, Garrett M. Morris, Kian Ming A. Chai
FDARxBench is an expert‑curated benchmark designed to evaluate document‑grounded question answering on FDA drug label documents, focusing on generic drug assessment. It features a multi‑stage pipeline that generates high‑quality QA examples covering factual, multi‑hop, and refusal tasks, and includes protocols for both open‑book and closed‑book reasoning. Experiments with various language models show significant gaps in factual grounding, long‑context retrieval, and safe refusal behavior, highlighting the challenge of regulatory‑grade label comprehension.
By Betty Xiong, Jillian Fisher, Benjamin Newman, Meng Hu, Shivangi Gupta, Yejin Choi, Lanyan Fang, Russ B Altman
ChemMLLM is a unified chemical multimodal large language model designed for molecule understanding and generation across text, SMILES strings, and images. The authors curated five multimodal tasks and benchmarked ChemMLLM against leading general MLLMs, chemical LLMs, and specialized models, finding it outperforms general-purpose MLLMs and matches specialized models on all tasks. The study demonstrates that a single foundation model can handle diverse cross‑modal chemical tasks, including image generation, enabling more intuitive visual human‑AI interaction.
By Qian Tan, Di Zhang, Ben Gao, Peng Xia, Wanhao Liu, Shufei Zhang, Wanli Ouyang, Lei Bai, Yuqiang Li, Tianfan Fu
arXiv:2510.26824v2 Announce Type: replace-cross
Abstract: Wide access to advanced experimental methods in materials science has given rise to an abundance of procedural knowledge, which is scattered...
By Magdalena Lederbauer, Siddharth Betala, Valerie Gentzke, Anamaria Leonescu, Amine Sehaba, Faris Flaifil, Ayush Jain, Alfonso Amayuelas, Nikhil Yelamarthy, Xiyao Li, Gr\'egoire Germain, Stefano Ribes, Stefan P. Schmid, Alexandre Nozadze, Anna Kelmanson, Sudheesh Kumar Ethirajan, Mohd Zaki, Elton Pan, Georgia Channing, Connor W. Coley, Philippe Schwaller, Roc\'io Mercado, Alexandre Duval, Mathilde L. D. Franckel, Samuel P. Gleason
arXiv:2606. 03660v1 Announce Type: new Abstract: Large language models are increasingly used as chemistry assistants, yet most chemistry benchmarks still score only final answers.
By Hongyu Guo, Hao Li, He Cao, Gongbo Zhang, Li Yuan
arXiv:2608. 11283v1 Announce Type: cross Abstract: Computation-ready metal-organic framework (MOF) databases are essential for high-throughput screening, yet many reported crystal structures remain chemically unreasonable or disordered, compromising simulation fidelity.
By Guobin Zhao, Xiao-Yan Li