The paper introduces Resolution-Aware Experimental Design (RAED), a method that selects experiments by minimizing the expected size of the nonempty structural candidate set while controlling false-exclusion rates. RAED is shown to preserve expected ordering under a composite Blackwell comparison and is implemented via a learned score-based approach with finite-sample nuisance-average and positive-tail calibration. Experiments on subsurface-flow, fluvial, and methane-oxidation benchmarks demonstrate RAED’s ability to resolve structural ambiguities and provide finite-sample guarantees for tail-sensitive nuisance risk.
By Sofianos Panagiotis Fotias
arXiv:2604. 11305v3 Announce Type: replace Abstract: Conformal selection (CS) uses calibration data to identify test inputs whose unobserved outcomes are likely to satisfy a pre-specified minimal quality requirement, while controlling the false discovery rate (FDR).
By Meiyi Zhu, Osvaldo Simeone
arXiv:2607. 11920v1 Announce Type: cross Abstract: Evaluating decisions made under uncertainty is hard when labeled outcomes are scarce, costly, or confounded with luck.
By Jeff Helzner
arXiv:2606. 04804v1 Announce Type: new Abstract: Generative models -- diffusion and flow matching -- are increasingly used to solve partial differential equation (PDE) inverse problems, enforcing the governing physics as a \emph{hard constraint} (via projection or guidance) and reporting the resulting samples as a Bayesian posterior with calibrated uncertainty.
By Jian Xu, Delu Zeng, John Paisley, Qibin Zhao
arXiv:2607. 16813v1 Announce Type: new Abstract: Sparse-support uncertainty is usually quantified by treating the dictionary as known, an assumption that can produce overconfident, label-dependent conclusions when the dictionary is learned from latent sparse mixtures.
By Guan-Ju Peng
arXiv:2608. 06262v1 Announce Type: new Abstract: Model evaluations may fix all tests before observing any responses or select later tests using earlier responses.
By Zonghuan Xu
arXiv:2606. 14690v1 Announce Type: new Abstract: We study a \emph{max-risk} objective for active learning in a multi-group mean estimation $d$-armed bandits: a learner adaptively allocates a budget of $T$ samples across $d$ groups to minimize the worst-case uncertainty index $\max_{k\in[d]}\sigma_k^2/n_k$, where $\sigma_k$ is the standard deviation of the distribution of arm $d$, and $n_k$ is the number of times arm $d$ is sampled.
By Abdellah Aznag, Rachel Cummings, Adam N. Elmachtoub
Generative models -- diffusion and flow matching -- are increasingly used to solve partial differential equation (PDE) inverse problems, enforcing the governing physics as a \emph{hard constraint} (via projection or guidance) and reporting the resulting samples as a Bayesian posterior with calibrated uncertainty. We show that this widely adopted recipe samples the wrong distribution.
arXiv:2607. 21721v3 Announce Type: replace-cross Abstract: Where truths are scarce (e.
By Ali Siahkoohi, Sina Alemohammad
arXiv:2608. 12652v1 Announce Type: cross Abstract: Benchmark contamination is diagnosed today with n-gram overlap, with likelihood-based membership inference, or with canary strings, and each needs something usually unavailable: the training corpus, a well-chosen test statistic, or foresight at dataset release.
By Florian Braun
arXiv:2606. 15877v1 Announce Type: cross Abstract: Chain-of-thought (CoT) improves large language models' performance in math and symbolic reasoning.
By Alex Bogdan
arXiv:2606. 15393v1 Announce Type: cross Abstract: Scientific discovery relies on large-scale hypothesis testing.
By Binyamin Perets, Shie Mannor