The paper introduces Resolution-Aware Experimental Design (RAED), a method that selects experiments by minimizing the expected size of the nonempty structural candidate set while controlling false exclusions. RAED is shown to align with a composite Blackwell comparison and is implemented via a learned score-based approach with finite-sample calibration. Experiments on subsurface-flow, fluvial, and methane-oxidation benchmarks demonstrate that RAED can diverge from expected-information-gain selections, yielding clearer resolution and explicit ambiguity handling.
arXiv:2604. 11305v3 Announce Type: replace Abstract: Conformal selection (CS) uses calibration data to identify test inputs whose unobserved outcomes are likely to satisfy a pre-specified minimal quality requirement, while controlling the false discovery rate (FDR).
By Meiyi Zhu, Osvaldo Simeone
arXiv:2607. 11920v1 Announce Type: cross Abstract: Evaluating decisions made under uncertainty is hard when labeled outcomes are scarce, costly, or confounded with luck.
By Jeff Helzner
arXiv:2607. 16813v1 Announce Type: new Abstract: Sparse-support uncertainty is usually quantified by treating the dictionary as known, an assumption that can produce overconfident, label-dependent conclusions when the dictionary is learned from latent sparse mixtures.
By Guan-Ju Peng
The paper investigates the safety of Code World Models, where a language model generates executable world models that a planner uses. It shows that accepting a model based on sampled transitions only guarantees sample consistency, not full safety, because the probability of missing critical events decays as (1‑r)^N. Experiments on hybrid instruments reveal that omitted mode‑boundaries can severely limit planner performance, and that even sophisticated LLMs (GPT‑5.x) struggle to repair such omissions in higher‑dimensional settings.
By Javier Aguilar Mart\'in
arXiv:2606. 04804v1 Announce Type: new Abstract: Generative models -- diffusion and flow matching -- are increasingly used to solve partial differential equation (PDE) inverse problems, enforcing the governing physics as a \emph{hard constraint} (via projection or guidance) and reporting the resulting samples as a Bayesian posterior with calibrated uncertainty.
By Jian Xu, Delu Zeng, John Paisley, Qibin Zhao
arXiv:2607. 05806v1 Announce Type: new Abstract: Training data for machine learning is routinely collected by a selection process the model never sees: loans are observed only when granted, outcomes only when a test was ordered.
By Gunner Levi Howe
arXiv:2606. 19361v1 Announce Type: cross Abstract: Identification conditions describe the computability of a target query or parameter of interest as a function of the type and amount of information available.
By Lucius E. J. Bynum, Rajesh Ranganath, Kyunghyun Cho
arXiv:2607. 21480v1 Announce Type: new Abstract: An initial high-recall stage in an empirical pipeline decides which items pass to later review, labelling, or modelling, and relevant items it misses are lost to every subsequent stage.
By Martin Anthony, Kaveh Salehzadeh Nobari
arXiv:2608. 06262v1 Announce Type: new Abstract: Model evaluations may fix all tests before observing any responses or select later tests using earlier responses.
By Zonghuan Xu
arXiv:2606. 14690v1 Announce Type: new Abstract: We study a \emph{max-risk} objective for active learning in a multi-group mean estimation $d$-armed bandits: a learner adaptively allocates a budget of $T$ samples across $d$ groups to minimize the worst-case uncertainty index $\max_{k\in[d]}\sigma_k^2/n_k$, where $\sigma_k$ is the standard deviation of the distribution of arm $d$, and $n_k$ is the number of times arm $d$ is sampled.
By Abdellah Aznag, Rachel Cummings, Adam N. Elmachtoub
arXiv:2608. 12652v1 Announce Type: cross Abstract: Benchmark contamination is diagnosed today with n-gram overlap, with likelihood-based membership inference, or with canary strings, and each needs something usually unavailable: the training corpus, a well-chosen test statistic, or foresight at dataset release.
By Florian Braun