The paper evaluates three AI model security scanners—ModelScan, ModelAudit, and Fickling—using a benchmark of 170 Pickle and PyTorch artifacts from 145 families, 135 of which have binary security labels. It distinguishes coverage metrics such as non‑N/A coverage, analysis completion, and definitive security decisions, finding that ModelAudit achieved 100% definitive decisions, Fickling 81.5%, and ModelScan 49.6%. When a definitive judgment was made, ModelScan reached perfect precision, recall, and F1, while Fickling added no unique true positives beyond those found by the other tools.
By Qianlong Lan, Vinothini Pandurangan, Anuj Kaul, Indranil Sanyal
The study audited ten different classifiers—including linear, tree‑ensemble, neural, glass‑box, and tabular foundation models—on national health survey data to predict myocardial infarction. By systematically removing features that could cause target leakage, the authors found that all models’ AUROC scores collapsed into a narrow band, indicating that reported high accuracy in prior work was largely due to leakage rather than model sophistication. The glass‑box explainable boosting machine performed comparably to other models while being much faster, and the authors demonstrated that fairness, calibration, and uncertainty can be audited and repaired without sacrificing performance.
By Raad Bin Tareaf, Murad Al-Rajab, Samia Loucif, Samer Ellaham, Cedric Schmitz
arXiv:2608. 16147v1 Announce Type: new Abstract: Class-imbalance handling is routinely evaluated on a single benchmark dataset, and the resulting conclusions are reported as if they were properties of the method.
By Diyorbek Musaev
arXiv:2608. 13601v1 Announce Type: new Abstract: Active learning can reduce labeling cost by selecting informative examples, but the most uncertain examples may also be the hardest to label correctly.
By John Myron Uy
arXiv:2606. 03305v1 Announce Type: new Abstract: Benchmark contamination, where evaluation examples appear in a model's training data, threatens the validity of LLM assessment.
By Wojciech Zarzecki, Jan Dubi\'nski, Sebastian Cygert
The paper introduces a new taxonomy for benchmark contamination that categorizes leakage by the mitigation it defeats—direct, derivative, temporal, distributional, and acquired—covering both training‑time and evaluation‑time scenarios. It proposes a four‑field disclosure protocol to record contamination status alongside benchmark scores, and provides a JSON schema, validator, and examples. An empirical study of 41 documents using a pre‑registered instrument shows limited reporting of contamination types and variable reliability, highlighting gaps in current disclosure practices.
By Johanna Angulo, V\'ictor Yeste, Hector Espinos-Morato