arXiv Machine Learning

PertReason: A Knowledge-Grounded Benchmark and Framework for Cell-State-Conditioned Mechanistic Reasoning of Perturbation Effects

arXiv:2607. 18777v1 Announce Type: new Abstract: Evaluating machine learning in scientific domains requires separating correct predictions from correct reasons under realistic distribution shifts.

arXiv AI
Aug 18

PertMind: Eliciting Emergent Biological Reasoning in LLM via Reinforcement Learning on Cellular Perturbation Data

arXiv:2608. 16419v1 Announce Type: cross Abstract: Large language models can describe mechanisms, yet scalable post-training still depends on costly, manually curated biological reasoning traces.

By Zhenchao Tang, Xiaogang Xu, Tianxu Lv, Jiahui Guan, Jiale Zhou, Haohuai He, Zhi Song, Hanbo Huang, Jiehui Huang, Jiafei Wu, Zhe Liu
arXiv AI
4d ago

OmniVCBench: Benchmarking Evidence-Grounded Multimodal Reasoning Towards AI Virtual Cells

OmniVCBench is a figure‑centric, source‑traceable benchmark designed to evaluate the interpretation component of Artificial Intelligence Virtual Cells (AIVCs). It comprises 6,077 curated question–answer pairs drawn from scientific figures and experimental contexts, organized into three scientific reasoning tasks that mirror the AIVC Predict–Explain–Discover agenda. The benchmark also introduces AIVC‑Judge, a task‑conditioned MLLM‑as‑a‑judge framework with reference‑aware rubrics, and a Model‑Derived Hard‑Negative Mining strategy to generate multiple‑choice distractors for efficient evaluation.

By Manyu Li, Xunkai Li, Yongfu Xiong, Yi Liu, Rong-Hua Li, Guoren Wang
arXiv Machine Learning
Sep 22

Tool-Augmented On-Policy Distillation for LLM Domain Adaptation in Sequence-Based Omics Tasks

The paper introduces OmicsBench, a new reasoning benchmark for multi‑omics sequences that includes 1,160 expert‑validated questions across DNA regulation, RNA processing, and protein function tasks, requiring traceable evidence chains. Evaluation of 17 large language models shows that scientific LLMs, while more accurate in classification, often lack valid evidence, suggesting shortcut learning. To address this, the authors propose tool‑augmented on‑policy distillation (TA‑OPD), a post‑training method that improves both evidence grounding and predictive performance across five Qwen3.5 models of varying sizes.

By Jie Ying, Zhefan Wang, Zihong Chen, Zhengqing Li, Jinzhe Li, Gang Li, Jian Liu, Fang Hu, Tao Luo, Zhonghang Yuan, Wanli Ouyang, Stan Z. Li, Fan Yang, Nanqing Dong
arXiv AI
Aug 13

Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence

arXiv:2608. 12036v1 Announce Type: new Abstract: AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their capabilities and the risks they may pose remain poorly understood.

By Mengru Wang, Junfeng Fang, Shuofei Qiao, Zhenqian Xu, Haoming Xu, Haoxiong Wang, Shumin Deng, Linyi Yang, Zhixiang Cui, Xin Xu, Yunzhi Yao, Buqiang Xu, Fei Shen, Haozhe Luo, Yunxiang Wei, Ningyu Zhang, Julian McAuley, Tat Seng Chua, Huajun Chen
arXiv Machine Learning
1d ago

When Do Biological Reasoning Models Use Their Biological Inputs?

The study evaluates whether biological reasoning models actually use their biological inputs by testing six models on DNA, protein, and single‑cell tasks. By perturbing one biological input while keeping others fixed, the authors find that many models (e.g., Evo2, ESM3, BioReason, BioReason‑Pro) rely primarily on textual information, with minimal impact from the biological representations. In contrast, models like ChatNT, Prot2Text‑V2, CellWhisperer, and Cell2Sentence‑Scale show greater dependence on their biological inputs, yet overall accuracy gains do not consistently reflect increased biological input contribution.

By Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik
arXiv AI
Jun 4

SCI-PRM: A Tool Aware Process Reward Model for Scientific Reasoning Verification

arXiv:2606. 04579v1 Announce Type: new Abstract: While Process Reward Models (PRMs) have achieved remarkable success in mathematical reasoning, their application in complex scientific domains-such as biology, chemistry, and physics remains largely unexplored.

By Xiangyu Zhao, Hengyuan Zhao, Yiheng Wang, Wanghan Xu, Yuhao Zhou, Qinglong Cao, Zhiwang Zhou, Lei Bai, Wenlong Zhang, Xiao-Ming Wu