arXiv:2511. 20851v3 Announce Type: replace-cross Abstract: Feature selection remains difficult in modern high-dimensional settings, and established methods such as Boruta and Recursive Feature Elimination are either computationally costly or lack a statistically justified stopping criterion for their importance scores.
By Mousam Sinha, Tirtha Sarathi Ghosh, Koushik Biswas, Ridam Pal
arXiv:2606. 01566v1 Announce Type: new Abstract: Small-to-medium scientific datasets place machine learning pipelines under two compounding pressures.
By Amanda S Barnard
arXiv:2412. 18134v5 Announce Type: replace Abstract: Randomized self-reductions (RSRs) express $f(x)$ using $f$ evaluated at random correlated points, enabling self-correcting programs, instance-hiding protocols, and applications in complexity theory and cryptography.
By Ferhat Erata, Orr Paradise, Thanos Typaldos, Timos Antonopoulos, ThanhVu Nguyen, Shafi Goldwasser, Ruzica Piskac
arXiv:2608.27704v1 Announce Type: new
Abstract: When machine learning classifiers are retrained, inputs correctly classified by the previous model version may be misclassified by the updated version,...
By Madhusudan Srinivasan, Namith Nishal Raphae
arXiv:2604. 11305v3 Announce Type: replace Abstract: Conformal selection (CS) uses calibration data to identify test inputs whose unobserved outcomes are likely to satisfy a pre-specified minimal quality requirement, while controlling the false discovery rate (FDR).
By Meiyi Zhu, Osvaldo Simeone
arXiv:2306. 14851v5 Announce Type: replace-cross Abstract: Given a high-dimensional covariate matrix and a response vector, ridge-regularized sparse linear regression selects a subset of features that explains the relationship between covariates and the response in an interpretable manner.
By Ryan Cory-Wright, Andr\'es G\'omez
arXiv:2606. 23880v1 Announce Type: new Abstract: From climate teleconnections to gene regulation, modern time-series datasets encompass tens or hundreds of interacting variables, making causal discovery increasingly challenging.
By Mohammad Fesanghary, Abhinav Havaldar
arXiv:2412. 12807v4 Announce Type: replace-cross Abstract: Selective classification is a powerful tool for automated decision-making in high-risk scenarios, allowing classifiers to act only when confident and abstain when uncertainty is high.
By Mohamed Ndaoud, Peter Radchenko, Bradley Rava
arXiv:2510. 14331v3 Announce Type: replace Abstract: We study program-learning methods that are efficient in both samples and computation.
By Shivam Singhal, Priyadarsi Mishra, Eran Malach, Tomer Galanti
arXiv:2605. 04954v2 Announce Type: replace-cross Abstract: Per-instance algorithm selection (PIAS) takes advantage of complementarity between a set of algorithms by deciding which algorithm to run on a given instance.
By Koen van der Blom, Diederick Vermetten
arXiv:2609.36307v1 Announce Type: new
Abstract: Partial Least Squares (PLS) regression extracts a few outcome-aligned directions in a high-dimensional X and is widely used across applied science, but...
By Pawe{\l} Lenartowicz, Hubert Plisiecki
MECHVAR is a lightweight, auditable rule for selecting experiments from a finite library to discriminate between candidate mechanisms. It chooses probes by maximizing the posterior‑weighted variance of predicted responses, a score that aligns with the Box–Hill pairwise‑KL criterion and links to expected information gain when separations are small. Experiments on a 25‑block audit and a Digits loop show MECHVAR outperforming confirmation‑first strategies and matching or exceeding EIG in identification accuracy while being far faster to compute.
By Yifan Guo