arXiv AI

MOF-VERIFY: A Failure-Aware Agentic Harness for MOF Hypothesis Verification

MOF-VERIFY is a failure-aware agentic harness designed to improve hypothesis verification for metal‑organic frameworks (MOFs). It introduces a diagnostic benchmark with four task families—structural grounding, synthesis‑condition verification, evidence‑sufficiency verification, and MLIP‑based computational verification—to pinpoint failures in knowledge access, evidence acquisition, and reasoning. Guided by these diagnostics, MOF‑Verify addresses structural, literature, evidence‑sufficiency, and computational bottlenecks, achieving significant performance gains over direct inference and retrieval‑based baselines across multiple large language models.

arXiv AI
Sep 18

oMeBench: Towards Robust Benchmarking of LLMs in Organic Mechanism Elucidation and Reasoning

oMeBench is a large-scale, expert-curated benchmark designed to evaluate large language models (LLMs) on organic mechanism reasoning. It contains over 10,000 annotated mechanistic steps, including reaction type labels, intermediate structures, and difficulty ratings, and introduces the oMeS scoring framework to assess logical consistency and chemical structural similarity. Evaluation shows that while current LLMs display promising chemical intuition, they often fail to produce correct and consistent multi-step reasoning, though prompting and fine-tuning can bring smaller models up to the level of closed‑source frontier models.

By Ruiling Xu, Yifan Zhang
arXiv AI
Aug 24

An LLM agent for end-to-end computational materials discovery

MAESTRO is a large language model agent that automates the full screening pipeline for metal‑organic frameworks (MOFs). It parses extensive MOF literature, links publications to crystal structures, curates a computation‑ready database, and then applies a progressively more expensive computational strategy to identify promising candidates. The identified materials for wet flue gas separation come from unrelated studies, demonstrating the agent’s ability to uncover high‑performance materials across domains.

By Chen Yuntong, Huang Ju, Liu Yu, Zhao Dan, Sun Mingqi, Ju Chentian, Liu Yanbing, Huang Lijiang, Zhao Guobin
arXiv Machine Learning
Jun 9

Enhancing Spatial Reasoning in Large Language Models for Metal-Organic Frameworks Structure Prediction

arXiv:2601. 09285v2 Announce Type: replace Abstract: Metal-organic frameworks (MOFs) are porous crystalline materials with broad applications such as carbon capture and drug delivery, yet accurately predicting their 3D structures remains a significant challenge.

By Mianzhi Pan, JianFei Li, Peishuo Liu, Botian Wang, Yawen Ouyang, Yiming Rong, Hao Zhou, Jianbing Zhang
arXiv AI
Sep 23

ChemVTS-Bench: Evaluating Visual-Textual-Symbolic Reasoning of Multimodal Large Language Models in Chemistry

ChemVTS-Bench is a domain-authentic benchmark that evaluates Visual‑Textual‑Symbolic reasoning in multimodal large language models for chemistry. It presents diverse chemical problems—organic molecules, inorganic materials, and 3D crystal structures—in three input modes: visual-only, visual‑text hybrid, and SMILES-based symbolic. The benchmark includes an automated agent workflow for inference, answer verification, and failure diagnosis, and shows that visual-only inputs and structural chemistry remain challenging for current models.

By Zhiyuan Huang, Baichuan Yang, Zikun He, Yanhong Wu, Fang Hongyu, Zhenhe Liu, Lin Dongsheng, Bing Su
arXiv AI
Aug 11

El Agente Gr\'afico: A Semantic Execution Runtime for Scientific Agents

arXiv:2602. 17902v2 Announce Type: replace Abstract: Large language models (LLMs) can plan scientific workflows and generate code, but these capabilities do not specify how scientific state is validated, transferred and recorded across heterogeneous computational and experimental operations.

By Jiaru Bai, Abdulrahman Aldossary, Thomas Swanick, Marcel M\"uller, Yeonghun Kang, Changhyeok Choi, Naruki Yoshikawa, Zijian Zhang, Jin Won Lee, Tsz Wai Ko, Aiwei Yin, Mohammad Ghazi Vakili, Chris Crebolder, Varinia Bernales, Al\'an Aspuru-Guzik
arXiv AI
Sep 12

Autonomous Chemical Mechanistic Discovery through Agentic Reasoning and Validation

The paper introduces ARCHE, an autonomous system that combines a general-purpose reasoning model, a domain-specialized computational chemistry model, and a structured tool registry to automate chemical mechanism discovery. ARCHE interprets scientific questions, generates and prioritizes mechanistic hypotheses, orchestrates computational workflows, and refines conclusions in a closed loop. The authors validate the system on three challenging scenarios, including reconstructing stereocontrolling transition states, proposing a radical pathway for an unpublished reaction, and identifying a descriptor governing selectivity in nickel-catalyzed cross‑coupling reactions.

By Dong Li, Sixuan Mi, Zihao Ye, Huan Xiong, Tao XU, Tong Zhu, Aijia Zhang, Junqi Gao, Kaiyan Zhang, Shijie Wang, Bowen Zhou, Yuqiang Li, Biqing Qi
arXiv AI
Sep 18

SCICONVBENCH: Benchmarking LLMs on Multi-Turn Clarification for Task Formulation in Computational Science

SCICONVBENCH is a benchmark designed to evaluate large language models (LLMs) on multi‑turn clarification tasks in computational science. It focuses on two key abilities: eliciting missing information (disambiguation) and resolving contradictory requests (inconsistency resolution) across four domains—fluid mechanics, solid mechanics, materials science, and partial differential equations. The benchmark pairs a structured task ontology with a rubric‑based evaluation framework, measuring LLM performance in clarification behavior, conversational grounding, and final‑specification fidelity, and reveals that even top models only resolve about 52.7% of disambiguation cases in fluid mechanics while often making ungrounded assumptions.

By Nithin Somasekharan, Youssef Hassan, Shiyao Lin, Gihan Panapitiya, Patrick Emami, Anurag Acharya, Sameera Horawalavithana, Shaowu Pan