EvoTreeNAD is a genealogy‑guided evolutionary algorithm that autonomously discovers neural architectures without a predefined seed or search space. Starting from an empty root, it builds a persistent genealogy where each node represents a complete architecture; top‑percentile values from nodes and descendants steer lineage selection. The method combines an Idea Agent that proposes variants and a Code Agent that implements them, with theoretical analysis showing stationary variation regimes and empirical results demonstrating superior performance on CIFAR‑10/100 and MedMNIST‑v2 tasks.
By Lishan Yu, Derek Jiu, Qizhen Lan, Xiaoqian Jiang
BioPhys-Bridge is a newly released benchmark dataset designed to evaluate language models on evidence‑grounded scientific reasoning within biophysical literature. Each of its 500 cases includes evidence blocks, stable IDs, quantitative values, units, equations, assumptions, mechanisms, and next‑step decisions, covering six biological domains and nine physical model families. The dataset enforces strict quality gates and has already been evaluated against several models, with DeepSeek‑V4‑Flash achieving the highest evidence‑ID F1 score of 0.360.
By Qingyang Xu
arXiv:2606. 01286v1 Announce Type: cross Abstract: The rapid progress of frontier large language models has led to widespread benchmark saturation, limiting the ability of existing datasets to differentiate model capabilities or provide useful training signal.
By Yangzhen Wu, Aaron J. Li, Wenjie Ma, Li Cao, Ziheng Zhou, Mert Cemri, Shu Liu, Yuran Xiu, Chenxiao Yan, Haikun Zhao, Bin Yu, Ion Stoica, Dawn Song
LongWoF-Bench is a new benchmark of 778 machine‑verifiable long‑workflow tasks spanning code generation, agent‑environment synthesis, mathematical reasoning, and rule following. The study shows that EvoMap Genes—structured representations of verifier‑confirmed execution trajectories—outperform the Skill baseline by 8.7–15.5 percentage points across seven models, and for Claude Opus they enable 39 additional task completions while cutting token consumption by 9.9%. The results demonstrate that verified execution experience can be externalized and reused, improving long‑workflow completion without repeatedly discovering new strategies.
By Xiao Zhang, Qumeng Sun, Jihao Li, Yiming Ren, Xiang Liu, Haoyang Zhang, Junjie Wang
Evolutionary Ensemble Search (EES) is a framework that builds machine‑learning procedures through expert‑guided program evolution. A specialized council interprets task evidence and experimental results to generate structured search directions, which an orchestrator assigns to execution specialists and an evolutionary engine. The engine selects parents, diagnoses errors, and creates descendants via code mutation, pipeline edits, and crossover, with each child evaluated on its own validation evidence. Population archives preserve useful alternatives, and compatible predictions compete in a validation‑gated ensemble stage. Search adapts through parent‑relative operator credit, session memory, and lessons retrieved across runs. The system achieved medal‑threshold artifacts on 19 of 22 tasks (86.36 %) with 11 gold, five silver, and three bronze outcomes across diverse modalities.
By Juan P. Madrigal-Cianci, Eshan Chordia
arXiv:2608. 03569v1 Announce Type: new Abstract: Benchmarking the ability of AI scientists to generate novel ideas is notoriously difficult.
By William Bolton, Philip Torr
arXiv:2607. 04439v1 Announce Type: new Abstract: Large language models have made research ideation increasingly accessible, yet effective idea development requires more than generating candidate directions.
By Qihao Zhao, Yangyu Huang, Yalun Dai, Lingao Xiao, Jianjun Gao, Xin Zhang, Wenshan Wu, Scarlett Li, Yang He, Yan Lu, Yap Kim Hui
Large language models are increasingly expected to execute complex workflows whose success depends on maintaining interdependent constraints and producing artifacts that satisfy strict end-to-end veri...
arXiv:2609.07611v1 Announce Type: new
Abstract: Scientific ideation is the capacity to formulate novel and testable hypotheses from scientific evidence, and autonomous AI scientists depend on it. Exi...
By Yunxiang Mo, Tianshi Zheng, Yisen Gao, Rui Wang, Newt Nguyen Kim Hue Nam, Kelvin Kiu Wai Tam, Jiaxin Bai, Yangqiu Song, Ginny Wong, Simon See
The paper introduces OmicsBench, a new reasoning benchmark for multi‑omics sequences that includes 1,160 expert‑validated questions across DNA regulation, RNA processing, and protein function tasks, requiring traceable evidence chains. Evaluation of 17 large language models shows that scientific LLMs, while more accurate in classification, often lack valid evidence, suggesting shortcut learning. To address this, the authors propose tool‑augmented on‑policy distillation (TA‑OPD), a post‑training method that improves both evidence grounding and predictive performance across five Qwen3.5 models of varying sizes.
By Jie Ying, Zhefan Wang, Zihong Chen, Zhengqing Li, Jinzhe Li, Gang Li, Jian Liu, Fang Hu, Tao Luo, Zhonghang Yuan, Wanli Ouyang, Stan Z. Li, Fan Yang, Nanqing Dong
arXiv:2604.15097v3 Announce Type: replace-cross
Abstract: This beta technical report asks how reusable experience should be represented so that it can function as effective test-time control and as a...
By Junjie Wang, Yiming Ren, Haoyang Zhang
arXiv:2608. 08958v1 Announce Type: cross Abstract: Tree Search-based test-time scaling of LLMs is a powerful tool for automated scientific coding.
By Xuefei Julie Wang, Hao Cui, Michael P. Brenner, Subhashini Venugopalan