arXiv:2606. 01566v1 Announce Type: new Abstract: Small-to-medium scientific datasets place machine learning pipelines under two compounding pressures.
By Amanda S Barnard
arXiv:2608.27704v1 Announce Type: new
Abstract: When machine learning classifiers are retrained, inputs correctly classified by the previous model version may be misclassified by the updated version,...
By Madhusudan Srinivasan, Namith Nishal Raphae
The paper investigates how making responsible‑AI evaluations more efficient—through batching, quantization, and benchmark reduction—affects the stability of conclusions drawn about model behavior. By testing three dense and mixture‑of‑experts models on the BBQ and BBQ‑V datasets under seven different conditions, the authors compare accuracy, bias, reasoning quality, subgroup performance, subset‑membership stability, runtime, and GPU energy consumption against a full‑benchmark BF16 baseline. Findings show that larger batching preserves accuracy and reduces energy in most settings, INT8 largely maintains quality but can increase energy use, INT4 introduces larger, context‑dependent changes, and reduced benchmarks save resources but are highly sensitive to which items are retained, underscoring that efficient evaluation must be validated against the benchmark’s intended conclusions.
By Ahmed El Kady, Aravind Narayanan, Rehana Noorani, Yani Ioannou, Shaina Raza
arXiv:2607. 07060v1 Announce Type: cross Abstract: Inherently interpretable classifiers for tabular data typically rely on sparse features, rules, or patterns that users can inspect directly.
By Srikumar Krishnamoorthy
arXiv:2606. 29280v1 Announce Type: cross Abstract: We identify intervention bias as a previously unquantified failure mode of zero-shot large-language-model (LLM) educational advisory agents: without task-specific training, they recommend action when a hindsight-optimal oracle policy mandates inaction.
By Craig Atkinson
The paper studies Online Kernel Supervised Principal Component Analysis (OKSPCA), which uses random features and an Adam-style orthonormal basis update to optimize a supervised spectral objective. It shows that accurate optimization of this objective does not guarantee accurate population subspace recovery or improved predictive performance, and it provides theoretical results on consistency, concentration, and perturbation of the estimator. Empirical experiments on six benchmarks reveal that replacing the tracker with the exact empirical target does not significantly change regression deficits, while classification-rank models capture most of the terminal objective energy but can exhibit substantial geometric deviation; sample-size studies further separate empirical accuracy from population recovery. The diagnostics also compare computational trade-offs, indicating that exact on-request computation can be faster in classification settings, whereas Adam saves time relative to full thin‑SVD in some dense regression requests, despite persistent geometric error.
By Zhenlin Yao, Wei Xiong
Efficient evaluation changes the protocol used to support claims about model behavior, yet it is rarely tested whether those claims remain stable after the evaluation itself is made cheaper. We stress...
Data-DPO is a target model‑oriented supervised fine‑tuning data selection method that uses one‑step probing of the target model to generate pairwise data preferences, trains a lightweight reward model to capture these preferences, and then selects a training subset by combining target‑model preference, external quality scores, and marginal diversity. Experiments on Vision‑Flan and LLaVA‑CoT demonstrate that Data‑DPO consistently outperforms existing data selection baselines across multiple data budgets and even surpasses full data training performance.
By Peng Sun, Yi Yang, Antong Zhang, Chunxiao Li, Yanbo Wang, Dianbo Liu, xin chen, Kai Yu, Lu Chen, Tianfan Fu
arXiv:2605. 29411v2 Announce Type: replace-cross Abstract: Under standard graphical assumptions, the Markov boundary of a target variable is the smallest set of features that renders every other feature redundant.
By Shu Wan, Abhinav Gorantla, Huan Liu, K. Sel\c{c}uk Candan
arXiv:2607. 18278v1 Announce Type: cross Abstract: Calibration is usually evaluated in aggregate, but the most dangerous failures are often local: predictions that remain highly confident despite being wrong.
By Filippo Cenacchi, Longbing Cao, Runze Yang
arXiv:2607. 11969v1 Announce Type: cross Abstract: Point-adjustment (PA), long the default scoring protocol in time-series anomaly detection (TSAD), was shown by Kim et al.
By Zongye Lyu
arXiv:2607. 20046v1 Announce Type: cross Abstract: With the widespread deployment of deep neural networks (DNNs) in safety-critical domains, reducing the cost of model validation under limited testing budgets has become increasingly important.
By Chunyu Liu, Mingyuan Li, Yang Li, Wenmin Li, Fei Gao, Tengfei Tu, Su-Juan Qin