arXiv Machine Learning

Finite Resources False Discovery Rate Control in Structured Hypothesis Spaces

arXiv:2606. 15393v1 Announce Type: cross Abstract: Scientific discovery relies on large-scale hypothesis testing.

arXiv AI
Sep 24

False-science induction in autonomous scientific discovery

The paper investigates how closed‑loop autonomous discovery systems can develop false‑science induction when physical objects and measurements are incorrectly paired. It demonstrates that such misbinding causes neural surrogates to learn spurious associations, diverting experimental effort toward low‑performing regions in both green fluorescent protein fitness and materials band‑gap prediction loops. The study shows that the coherence of these errors—not just their frequency—drives budget misallocation and proposes monitoring strategies to detect and quarantine corrupted hypothesis axes.

By Hanbing Liang, Fujun Liu
Hugging Face Trending Papers
Sep 3

Resolution-Aware Experimental Design under Partial Identifiability

The paper introduces Resolution-Aware Experimental Design (RAED), a method that selects experiments by minimizing the expected size of the nonempty structural candidate set while controlling false exclusions. RAED is shown to align with a composite Blackwell comparison and is implemented via a learned score-based approach with finite-sample calibration. Experiments on subsurface-flow, fluvial, and methane-oxidation benchmarks demonstrate that RAED can diverge from expected-information-gain selections, yielding clearer resolution and explicit ambiguity handling.

arXiv Machine Learning
Sep 4

Resolution-Aware Experimental Design under Partial Identifiability

The paper introduces Resolution-Aware Experimental Design (RAED), a method that selects experiments by minimizing the expected size of the nonempty structural candidate set while controlling false-exclusion rates. RAED is shown to preserve expected ordering under a composite Blackwell comparison and is implemented via a learned score-based approach with finite-sample nuisance-average and positive-tail calibration. Experiments on subsurface-flow, fluvial, and methane-oxidation benchmarks demonstrate RAED’s ability to resolve structural ambiguities and provide finite-sample guarantees for tail-sensitive nuisance risk.

By Sofianos Panagiotis Fotias