arXiv Computation and Language By Liangcai Su, Zhaopeng Feng, Zhuo Chen, Zhen Zhang, Xiang Lin, Ruilin Li, Handuo Zhang, Ning Wang, Kailong Wen, Yueqi Guo, Feng Xing, Yiling Guo, Chenxiong Qian, Simon Shaolei Du, Lidong Bing, Xinyu Wang

FrontierChallenge: Evaluating Scientific Workflow Completion

Read the original on arXiv Computation and Language →

FrontierChallenge is a cross‑domain benchmark that releases 300 end‑to‑end scientific workflows, of which 97 are evaluated in this study. The benchmark covers diverse fields such as quantum chemistry, molecular dynamics, materials characterization, analytical chemistry, life science, and electrochemistry/environment, and requires agents to produce a bundle of fixed scientific deliverables. Twelve frontier models were tested, and the best configurations completed only 20 of the 97 tasks, achieving a 20.6% pass rate; high partial scores and confident completion claims often did not translate into full delivery, especially in analytical chemistry and electrochemistry/environment.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Sep 15

OpenAl4S: Code as Action, Science as Sessions

OpenAI4S is an open‑source scientific research agent that treats code as action and science as sessions, combining a persistent computing runtime with structured session management. It uses tool calls for orchestration, executes code cells in persistent Python and R kernels, and records an append‑only Action Ledger, per‑cell execution logs, versioned artifacts, environment snapshots, and workspace checkpoints to preserve provenance and enable session recovery, branching, and extension. Evaluated on 36 research scenarios—including retrosynthesis, molecular dynamics, and protein design—OpenAI4S achieved a higher overall score (7.83) than a general‑purpose coding harness, especially on long‑horizon, computation‑intensive workflows, though reproducibility remains an open challenge. whyItMatters":"The system demonstrates that persistent execution coupled with session‑level provenance can enhance the reliability of AI‑assisted scientific workflows, as evidenced by its superior performance across diverse research scenarios."

By Gongbo Zhang, Hao Li, Yu Wang, Mujie Lin, Liuzhenghao Lv, Yicheng Mao, Yimi Wang, Jun Zhu, Minhan Tang, Zhengxiang Jiang, Yusong Wang, Jiayu Yao, Kunpeng Ning, Dawei Pang, Yonghong Tian, OpenAI4S Community, Yuyang Liu, Li Yuan
arXiv AI
Aug 11

El Agente Gr\'afico: A Semantic Execution Runtime for Scientific Agents

arXiv:2602. 17902v2 Announce Type: replace Abstract: Large language models (LLMs) can plan scientific workflows and generate code, but these capabilities do not specify how scientific state is validated, transferred and recorded across heterogeneous computational and experimental operations.

By Jiaru Bai, Abdulrahman Aldossary, Thomas Swanick, Marcel M\"uller, Yeonghun Kang, Changhyeok Choi, Naruki Yoshikawa, Zijian Zhang, Jin Won Lee, Tsz Wai Ko, Aiwei Yin, Mohammad Ghazi Vakili, Chris Crebolder, Varinia Bernales, Al\'an Aspuru-Guzik