arXiv AI

Closed-Loop Molecular Design with Calibrated Deference

arXiv:2606. 02618v1 Announce Type: cross Abstract: We present Cognitive Loop via In-Situ Optimization (CLIO), an agent that couples a continuously-updated belief-state graph with a recursive plan-then-act loop.

arXiv AI
Sep 2

Autonomous discovery of new structure-plausibility laws for explainable and rapid crystal diagnosis and screening

The paper reports that autonomous agents can generate, test, and refute millions of candidate structure-plausibility laws, ultimately producing eight Plausibility Rules for Inorganic Structures (PRIS) that capture five key mechanisms of crystal stability. PRIS achieves high agreement with experimental structures (82–99%) and outperforms traditional Pauling rules, detecting damaged crystals with 87.9% accuracy and correlating strongly with synthesizability. By integrating PRIS with a synthesis score (PSS), the authors demonstrate significant reductions in expensive DFT validation and improved inverse-design efficiency, while also providing chemical explanations for failures and anomalies in crystal data.

By Zhilong Song, Lixue Cheng
arXiv AI
Jun 2

Probe Before You Edit: Probing-Guided Molecular Optimization for LLM Agents in Structure-Based Drug Design

arXiv:2606. 00555v1 Announce Type: new Abstract: Structure-based drug design increasingly employs LLM agents to iteratively refine ligands against a target pocket, yet a viable ligand must satisfy two often-conflicting objectives -- binding affinity and druggability -- which single optimization steps rarely improve together.

By Zaifei Yang, Weiyu Chen, Yaqing Wang, James Kwok
arXiv AI
Aug 5

MDArena: Evaluating Coding Agents on Realistic Molecular Dynamics Workflows

arXiv:2608. 02642v1 Announce Type: cross Abstract: Accelerating scientific discovery is among the most consequential applications of AI, and computational biomolecular simulation stands out as a particularly promising target within this broader effort.

By Nithishwer Mouroug Anand, Wei-Tse Hsu, Kyle Vaccaro, Eden James Gage, Jonathan David Colburn, Linda Xi Phan, Minjoon Seo, Kevin Guan, Philip C. Biggin
arXiv AI
Sep 12

Autonomous Chemical Mechanistic Discovery through Agentic Reasoning and Validation

The paper introduces ARCHE, an autonomous system that combines a general-purpose reasoning model, a domain-specialized computational chemistry model, and a structured tool registry to automate chemical mechanism discovery. ARCHE interprets scientific questions, generates and prioritizes mechanistic hypotheses, orchestrates computational workflows, and refines conclusions in a closed loop. The authors validate the system on three challenging scenarios, including reconstructing stereocontrolling transition states, proposing a radical pathway for an unpublished reaction, and identifying a descriptor governing selectivity in nickel-catalyzed cross‑coupling reactions.

By Dong Li, Sixuan Mi, Zihao Ye, Huan Xiong, Tao XU, Tong Zhu, Aijia Zhang, Junqi Gao, Kaiyan Zhang, Shijie Wang, Bowen Zhou, Yuqiang Li, Biqing Qi
arXiv AI
4d ago

R-GroundBench: A Diagnostic Benchmark for R-Group Groundingin Markush Molecular Editing

R-GroundBench is a new diagnostic benchmark for evaluating AI models on R‑group grounding in Markush molecular editing, derived from real pharmaceutical patents. It includes a Multiple‑Choice VQA track with varying difficulty and modality splits, as well as an open‑ended Generation track. Experiments show a large performance gap: models score over 90% on easy VQA but drop to 56–66% on hard VQA, and generation exact match stays below 20% (and under 8% with visual input).

By Xin Wang, Zichuan Ying, Xinna Lin, Junqi Zhang, Hanyi Xiong, Tianyu Gao, Hairong Zhang, Qixiang Hua, Botian Shi, Zhenhailong Wang, Kaicheng Yu
arXiv Computation and Language
Sep 15

Chemical Chain-of-Thought Functions as a Hallucination-Prone Molecular Scratchpad

Chemical reasoning language models are expected to produce faithful chain-of-thought (CoT) explanations when answering chemistry tasks, but across four model families and twelve tasks, hallucinations are widespread and largely independent of answer correctness. Attribution analyses reveal that these models use a shared scratchpad function: Chem‑R and ether‑0 rely on fragmented SMILES drafts, while ChemDFM‑R emphasizes scaffold, positional, and naming cues. Perturbing Chem‑R’s SMILES sketches degrades generation, indicating that structural drafts can be causally load‑bearing even when verbal structural claims are largely inert.

By Jiatong Li, Yuxuan Ren, Weida Wang, Xiaoyong Wei, Yatao Bian
arXiv AI
Sep 18

oMeBench: Towards Robust Benchmarking of LLMs in Organic Mechanism Elucidation and Reasoning

oMeBench is a large-scale, expert-curated benchmark designed to evaluate large language models (LLMs) on organic mechanism reasoning. It contains over 10,000 annotated mechanistic steps, including reaction type labels, intermediate structures, and difficulty ratings, and introduces the oMeS scoring framework to assess logical consistency and chemical structural similarity. Evaluation shows that while current LLMs display promising chemical intuition, they often fail to produce correct and consistent multi-step reasoning, though prompting and fine-tuning can bring smaller models up to the level of closed‑source frontier models.

By Ruiling Xu, Yifan Zhang
arXiv AI
Aug 10

Strategy-first synthesis planning for complex natural products

arXiv:2608. 07454v1 Announce Type: cross Abstract: The total synthesis of a complex molecule is among the most demanding intellectual and experimental feats in chemistry: a chemist must plan many steps ahead for how to assemble simple building blocks into an intricate target, devise backup strategies, and anticipate procedural challenges.

By Daniel Armstrong, Xuan-Vu Nguyen, Octavian Susanu, Gabriel Gibberd, Th\'eo A. Neukomm, Tadd\"aus Strunden, Dan Forster, Morgane Delattre, Shawn Teh, Cl\'ement Rols, John Federice, Hayden Leatherwood, M. Lavelle Barnes, Maarten R. Dobbelaere, Peter Wipf, Jon T. Njardarson, Jieping Zhu, Philippe Schwaller