An Agentic AI Framework to Accelerate Scientific Discovery in Plant Phenotyping
arXiv:2606. 31831v1 Announce Type: new Abstract: High-throughput plant phenotyping now generates image derived datasets far faster than scientists can analyze them.
High-throughput plant phenotyping now generates image derived datasets far faster than scientists can analyze them. At Oak Ridge National Laboratory's Advanced Plant Phenotyping Laboratory (APPL), automated stations image hundreds of plants daily across multiple remote sensing modalities; yet, trait extraction and interpretation remain manual, expert-bound, and strictly post-hoc, making analysis, not acquisition, the binding constraint on discovery.
arXiv:2606. 31831v1 Announce Type: new Abstract: High-throughput plant phenotyping now generates image derived datasets far faster than scientists can analyze them.
arXiv:2606. 12736v1 Announce Type: new Abstract: AI agents are increasingly being developed to accelerate scientific discovery, yet their practical capabilities in real research settings remain poorly understood.
arXiv:2607. 02771v1 Announce Type: new Abstract: Leadership computing facilities steward large-scale scientific datasets that routinely require substantial transformation before serving as AI training data.
arXiv:2607. 16845v1 Announce Type: new Abstract: Scientists at European XFEL conduct experiments that generate very large and complex datasets.
arXiv:2603. 01421v3 Announce Type: replace Abstract: While large language models accelerate scientific discovery, existing agents face severe limitations in adaptability, domain generalization, and multimodal scalability, often struggling to autonomously process raw, domain-specific experimental data.
arXiv:2606. 13662v1 Announce Type: new Abstract: LLM-based agents have shown increasing potential in automating scientific discovery.
The paper introduces ScientistTwo, a fully autonomous multi‑agent framework that takes a scientific problem, establishes baselines, generates hypotheses, and coordinates specialized agents to conduct an end‑to‑end discovery cycle without human intervention. It rigorously tests and refines its methods through automated experiments, ablation studies, and a closed‑loop peer‑review engine. Benchmarking against top conferences (ICLR, ICML, NeurIPS) shows that ScientistTwo produces expert‑level, publishable papers and codebases that outperform human state‑of‑the‑art models and receive higher review ratings under automated AI review.
The paper introduces the Scientific Data Skill (SciDSK), an agent‑ready representation that packages dataset‑specific knowledge and operational guidance as a reusable skill. SciDSK integrates dataset descriptions, scientific context, file organization, usage procedures, quality checks, and provenance information while keeping the data in its original repository. The authors define a structured specification, build a construction pipeline, and launch the Scientific Data Skill Bank to publish SciDSK resources across six scientific disciplines, demonstrating improved agent‑driven dataset discovery and interpretation through evaluation benchmarks.
arXiv:2608.22045v1 Announce Type: cross Abstract: Modern scientific facilities and instruments generate datasets at scales that are difficult for individual researchers to discover, access, and explo...
arXiv:2606. 02080v1 Announce Type: cross Abstract: Biological image analysis increasingly demands integration across heterogeneous tools, programming environments, and domain knowledge that few researchers can command simultaneously.
arXiv:2608. 19625v1 Announce Type: new Abstract: Scientific data are increasingly used by AI agents, yet existing dataset representations provide limited support for autonomous discovery, interpretation, and invocation.
arXiv:2605. 26305v2 Announce Type: replace Abstract: This paper details two novel frameworks for developing autonomous, agentic AI in scientific workflows.