arXiv:2606. 03305v1 Announce Type: new Abstract: Benchmark contamination, where evaluation examples appear in a model's training data, threatens the validity of LLM assessment.
By Wojciech Zarzecki, Jan Dubi\'nski, Sebastian Cygert
arXiv:2604. 08870v3 Announce Type: replace-cross Abstract: Student dropout is a persistent concern in Learning Analytics, yet comparative studies frequently evaluate predictive models under heterogeneous protocols, prioritizing discrimination over temporal interpretability and calibration.
By Rafael da Silva, Jeff Eicher, Gregory Longo
arXiv:2607. 10633v1 Announce Type: cross Abstract: Explainable machine learning (XML) pipelines applied to composite mental health outcomes can produce apparently-robust, cross-population-stable risk hierarchies that are largely artefacts of how the outcome was constructed.
By Alireza Dehghan, Negin Ashrafi
The paper introduces Counterfactual Fragility Certificates (CFC), a model‑agnostic audit protocol that maps each prediction to an evidence‑failure trajectory, summarizing it with metrics such as greedy flip budget, margin‑collapse area, degradation thresholds, and fragility dominance score. CFC is shown to identify brittle high‑confidence predictions on seven tabular benchmarks with an AUROC of 0.915, outperforming existing scalar scores by up to +0.405. The method remains effective across various perturbation and review‑budget scenarios, and can also inform fragility‑aware regularization and temperature correction.
By Filippo Cenacchi, Longbing Cao, Runze Yang
arXiv:2607. 16122v1 Announce Type: new Abstract: Evaluations should do more than measure a models current performance.
By Vipul Gupta, Zihao Wang, Razvan-Gabriel Dumitru, MohammadHossein Rezaei, Aakash Sabharwal, Yunzhong He
arXiv:2608. 00794v2 Announce Type: replace Abstract: Agentic AI evaluation pipelines produce benchmark scores that justify deployment decisions, safety certifications, and regulatory compliance claims.
By William Caban
arXiv:2606. 28471v1 Announce Type: new Abstract: Model capability is the central variable in LLM pre-training, yet is never observed directly: data shapes it prospectively, while evaluation reveals it only retrospectively, compressing samples, prompts, decoding, and scoring rules into one noisy score.
By Zhixuan Li, Jiangan Yuan, Han Xu
The paper presents a risk‑adaptive, evidence‑constrained framework for providing feedback in introductory programming courses. Using data from 2,993 failed submissions by 215 students, the authors built models that predict persistent failure and generate four tailored feedback conditions for 136 cases. The framework employs calibrated risk to decide when to intervene, evidence gating to limit feedback content, and a progressive assistance strategy that moves from self‑checks to localized hints.
By Shihao Wang
arXiv:2606. 01566v1 Announce Type: new Abstract: Small-to-medium scientific datasets place machine learning pipelines under two compounding pressures.
By Amanda S Barnard
The paper investigates how AI systems that perform best‑of‑n search require different validation strategies as the search width changes. It shows that auditing only small search widths leaves a gap in reliability estimates for larger widths, and proposes retaining candidate ranks and truth labels to estimate reliability across all widths up to N. The authors derive theoretical bounds on the minimax mean‑squared error, design procedures that achieve these bounds, and demonstrate that a shared audit can significantly reduce maximum error across many widths in practical CodeRM pools.
By Ricardo Fitas
The study explores whether combining traditional and digital learning analytics can predict failure in a first‑year CS1 course. Using data from 284 students across four cohorts, the authors identified ten candidate factors and built a logistic regression model that achieved 74.7% accuracy and 0.742 macro F1, with 87% recall for failing students. Weighted academic momentum, basic demographics, and LMS activity emerged as the most predictive features, suggesting that simple digital markers can enable early‑warning systems by week five.
By Lighton Phiri, Mutune Chaibela, Ivy Chisha, David Pungwa, Danny Siabbaba, Bydon Simukoko
The paper presents a risk‑adaptive, evidence‑constrained framework that uses learning analytics to provide personalized feedback in introductory programming. By training models on 2,993 failed‑submission states from 215 students, the authors predict persistent failure and generate four tailored feedback conditions for 136 cases. A calibrated risk policy selects interventions for 17.8% of eligible states, capturing 25.2% of persistent failures, and the framework ensures that generated messages contain all required components after evidence gating.