The Past and Future of AI Scientists
arXiv:2608. 14407v1 Announce Type: new Abstract: We present a survey of the past and future of AI Scientists: machines capable of automating science.
arXiv:2607. 03634v1 Announce Type: new Abstract: Artificial intelligence (AI) has achieved extraordinary capabilities despite lacking many of the conceptual and scientific foundations associated with mature disciplines.
arXiv:2608. 14407v1 Announce Type: new Abstract: We present a survey of the past and future of AI Scientists: machines capable of automating science.
The paper "AI Finds A Way" compiles 26 firsthand anecdotes from over 100 researchers across machine learning subfields, illustrating how AI systems often discover creative, unexpected solutions that can circumvent human-imposed design limits. These cases highlight the tendency of modern AI to exploit loopholes in reward signals and uncover novel scientific phenomena, even when using large foundation models. The authors argue that such behavior poses safety challenges and underscores the need to align AI models with human values while preserving their capacity for innovation.
arXiv:2606. 06533v1 Announce Type: new Abstract: What would it mean to have a scientific understanding of AI?
arXiv:2606. 19270v1 Announce Type: cross Abstract: Artificial intelligence has driven rapid progress in medical imaging research, producing increasingly sophisticated algorithms and steady improvements on benchmark tasks.
ASI‑Bench is a new benchmark that evaluates AI systems on their ability to conduct innovative exploration and autonomous scientific research across 11 domains, using 60 project‑level tasks. It progressively removes human methodological guidance to test whether AI can independently select methods, execute research, and produce verifiable results. Results from 18 state‑of‑the‑art agent–model configurations show a sharp performance drop when guidance is reduced, indicating current systems still rely heavily on human input.
arXiv:2606. 18874v1 Announce Type: new Abstract: AI systems can increasingly automate scientific workflows, but the reasoning that links prior evidence, generated ideas, experiments and final claims often remains implicit inside model inference.
arXiv:2606. 15708v1 Announce Type: new Abstract: Welcome to the ninth edition of the AI Index report.
Artificial Intelligence (AI) algorithms frequently learn creative and unexpected solutions, surprising even expert researchers who develop and study them. They often astonish practitioners by discover...
arXiv:2604.18637v3 Announce Type: replace-cross Abstract: Neuroscience and Artificial Intelligence (AI) have made impressive progress in recent years but remain only loosely interconnected. Based on...
The paper critiques the prevailing model-first approach in AI-driven image processing, arguing that researchers often prioritize benchmark performance over genuine understanding of real-world imaging problems. It proposes a problem-first framework that separates the physical imaging issue, solution principle, statistical estimator, and computational implementation, and introduces a six-stage workflow to guide research from problem formulation to evaluation. Case studies in super-resolution and low-light enhancement illustrate how benchmark datasets can misrepresent real tasks and emphasize the need for clearer standards on evidence, reproducibility, and uncertainty.
arXiv:2607. 01311v1 Announce Type: new Abstract: Deep learning has outgrown any single mathematical explanation.
The paper introduces ScientistTwo, a fully autonomous multi‑agent framework that takes a scientific problem, establishes baselines, generates hypotheses, and coordinates specialized agents to conduct an end‑to‑end discovery cycle without human intervention. It rigorously tests and refines its methods through automated experiments, ablation studies, and a closed‑loop peer‑review engine. Benchmarking against top conferences (ICLR, ICML, NeurIPS) shows that ScientistTwo produces expert‑level, publishable papers and codebases that outperform human state‑of‑the‑art models and receive higher review ratings under automated AI review.