arXiv AI By Yifan Zhou, Qihao Yang, Yan Li, Donggang Li, Xiru Hu, Hokin Deng, Ziyang Gong, Xuanyi Zhou, Huacan Wang, Xiangchao Yan, Wanghan Xu, Wenlong Zhang, Shaofeng Zhang, Yue Zhou, Yifan Yang, Zhihang Zhong, Xue Yang

Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation

Read the original on arXiv AI →

arXiv:2607. 08758v1 Announce Type: new Abstract: Scientific ideas rarely start from a blank page.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 25

EvoTreeNAD: Genealogy-Guided Evolution for LLM-Driven Neural Architecture Discovery

EvoTreeNAD is a genealogy‑guided evolutionary algorithm that autonomously discovers neural architectures without a predefined seed or search space. Starting from an empty root, it builds a persistent genealogy where each node represents a complete architecture; top‑percentile values from nodes and descendants steer lineage selection. The method combines an Idea Agent that proposes variants and a Code Agent that implements them, with theoretical analysis showing stationary variation regimes and empirical results demonstrating superior performance on CIFAR‑10/100 and MedMNIST‑v2 tasks.

By Lishan Yu, Derek Jiu, Qizhen Lan, Xiaoqian Jiang
arXiv AI
Sep 18

BioPhys-Bridge: A Benchmark for Interdisciplinary Scientific Reasoning in Physics-Grounded Biological Research

BioPhys-Bridge is a newly released benchmark dataset designed to evaluate language models on evidence‑grounded scientific reasoning within biophysical literature. Each of its 500 cases includes evidence blocks, stable IDs, quantitative values, units, equations, assumptions, mechanisms, and next‑step decisions, covering six biological domains and nine physical model families. The dataset enforces strict quality gates and has already been evaluated against several models, with DeepSeek‑V4‑Flash achieving the highest evidence‑ID F1 score of 0.360.

By Qingyang Xu
arXiv AI
Jun 2

BenchEvolver: Frontier Task Synthesis via Solution-Centric Evolution

arXiv:2606. 01286v1 Announce Type: cross Abstract: The rapid progress of frontier large language models has led to widespread benchmark saturation, limiting the ability of existing datasets to differentiate model capabilities or provide useful training signal.

By Yangzhen Wu, Aaron J. Li, Wenjie Ma, Li Cao, Ziheng Zhou, Mert Cemri, Shu Liu, Yuran Xiu, Chenxiao Yan, Haikun Zhao, Bin Yu, Ion Stoica, Dawn Song
arXiv Computation and Language
Aug 25

LongWoF-Bench: Evaluating EvoMap Genes for Verifiable Long-Workflow Tasks

LongWoF-Bench is a new benchmark of 778 machine‑verifiable long‑workflow tasks spanning code generation, agent‑environment synthesis, mathematical reasoning, and rule following. The study shows that EvoMap Genes—structured representations of verifier‑confirmed execution trajectories—outperform the Skill baseline by 8.7–15.5 percentage points across seven models, and for Claude Opus they enable 39 additional task completions while cutting token consumption by 9.9%. The results demonstrate that verified execution experience can be externalized and reused, improving long‑workflow completion without repeatedly discovering new strategies.

By Xiao Zhang, Qumeng Sun, Jihao Li, Yiming Ren, Xiang Liu, Haoyang Zhang, Junjie Wang
arXiv AI
Sep 17

Evolutionary Ensemble Search: Council-Guided Program Evolution with Persistent Memory

Evolutionary Ensemble Search (EES) is a framework that builds machine‑learning procedures through expert‑guided program evolution. A specialized council interprets task evidence and experimental results to generate structured search directions, which an orchestrator assigns to execution specialists and an evolutionary engine. The engine selects parents, diagnoses errors, and creates descendants via code mutation, pipeline edits, and crossover, with each child evaluated on its own validation evidence. Population archives preserve useful alternatives, and compatible predictions compete in a validation‑gated ensemble stage. Search adapts through parent‑relative operator credit, session memory, and lessons retrieved across runs. The system achieved medal‑threshold artifacts on 19 of 22 tasks (86.36 %) with 11 gold, five silver, and three bronze outcomes across diverse modalities.

By Juan P. Madrigal-Cianci, Eshan Chordia