arXiv Computation and Language By Jiatong Li, Yuxuan Ren, Weida Wang, Xiaoyong Wei, Yatao Bian

Chemical Chain-of-Thought Functions as a Hallucination-Prone Molecular Scratchpad

Read the original on arXiv Computation and Language →

Chemical reasoning language models are expected to produce faithful chain-of-thought (CoT) explanations when answering chemistry tasks, but across four model families and twelve tasks, hallucinations are widespread and largely independent of answer correctness. Attribution analyses reveal that these models use a shared scratchpad function: Chem‑R and ether‑0 rely on fragmented SMILES drafts, while ChemDFM‑R emphasizes scaffold, positional, and naming cues. Perturbing Chem‑R’s SMILES sketches degrades generation, indicating that structural drafts can be causally load‑bearing even when verbal structural claims are largely inert.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Jun 3

LatentChem: From Textual CoT to Latent Thinking in Chemical Reasoning

arXiv:2602. 07075v5 Announce Type: replace-cross Abstract: Current chemical large language models (LLMs) predominantly rely on explicit Chain-of-Thought (CoT) to solve complex reasoning problems.

By Xinwu Ye, Yicheng Mao, Yuxuan Liao, Jia Zhang, Yimeng Liu, Li Hao, Fang Wu, Zhiwei Li, Zehong Wang, Zhiyuan Liu, Zhenfei Yin, Li Yuan, Philip Torr, Huan Sun, xiangxiang Zeng, Mengdi Wang, Le Cong, Shenghua Gao, Xiangru Tang
arXiv AI
Sep 18

oMeBench: Towards Robust Benchmarking of LLMs in Organic Mechanism Elucidation and Reasoning

oMeBench is a large-scale, expert-curated benchmark designed to evaluate large language models (LLMs) on organic mechanism reasoning. It contains over 10,000 annotated mechanistic steps, including reaction type labels, intermediate structures, and difficulty ratings, and introduces the oMeS scoring framework to assess logical consistency and chemical structural similarity. Evaluation shows that while current LLMs display promising chemical intuition, they often fail to produce correct and consistent multi-step reasoning, though prompting and fine-tuning can bring smaller models up to the level of closed‑source frontier models.

By Ruiling Xu, Yifan Zhang
arXiv AI
Sep 23

ChemVTS-Bench: Evaluating Visual-Textual-Symbolic Reasoning of Multimodal Large Language Models in Chemistry

ChemVTS-Bench is a domain-authentic benchmark that evaluates Visual‑Textual‑Symbolic reasoning in multimodal large language models for chemistry. It presents diverse chemical problems—organic molecules, inorganic materials, and 3D crystal structures—in three input modes: visual-only, visual‑text hybrid, and SMILES-based symbolic. The benchmark includes an automated agent workflow for inference, answer verification, and failure diagnosis, and shows that visual-only inputs and structural chemistry remain challenging for current models.

By Zhiyuan Huang, Baichuan Yang, Zikun He, Yanhong Wu, Fang Hongyu, Zhenhe Liu, Lin Dongsheng, Bing Su