arXiv Machine Learning

Limited Structural Reliability in Public Educational Prediction Benchmarks: A Four-Dimension Audit of Seven Datasets

The study audited seven public educational prediction datasets using four pre‑modeling reliability checks—baseline gap, split instability, null separation, and metadata adequacy under group‑aware holdout. Only three datasets passed all checks; the others failed either group‑aware generalization tests or lacked necessary provenance metadata. The audit revealed that cross‑group fragility, rather than weak iid performance, was the dominant failure mode, and that increasing model complexity did not resolve these structural issues.

arXiv AI
Sep 2

Counterfactual Fragility Certificates: Exposing High-Confidence Brittleness under Structured Evidence Failure

The paper introduces Counterfactual Fragility Certificates (CFC), a model‑agnostic audit protocol that maps each prediction to an evidence‑failure trajectory, summarizing it with metrics such as greedy flip budget, margin‑collapse area, degradation thresholds, and fragility dominance score. CFC is shown to identify brittle high‑confidence predictions on seven tabular benchmarks with an AUROC of 0.915, outperforming existing scalar scores by up to +0.405. The method remains effective across various perturbation and review‑budget scenarios, and can also inform fragility‑aware regularization and temperature correction.

By Filippo Cenacchi, Longbing Cao, Runze Yang
arXiv AI
Sep 25

A Risk-Adaptive and Evidence-Constrained Framework for Generative AI Feedback in Programming Education

The paper presents a risk‑adaptive, evidence‑constrained framework for providing feedback in introductory programming courses. Using data from 2,993 failed submissions by 215 students, the authors built models that predict persistent failure and generate four tailored feedback conditions for 136 cases. The framework employs calibrated risk to decide when to intervene, evidence gating to limit feedback content, and a progressive assistance strategy that moves from self‑checks to localized hints.

By Shihao Wang
arXiv Machine Learning
Sep 15

The geometry of AI validation: From structural blindness to reusable audits

The paper investigates how AI systems that perform best‑of‑n search require different validation strategies as the search width changes. It shows that auditing only small search widths leaves a gap in reliability estimates for larger widths, and proposes retaining candidate ranks and truth labels to estimate reliability across all widths up to N. The authors derive theoretical bounds on the minimax mean‑squared error, design procedures that achieve these bounds, and demonstrate that a shared audit can significantly reduce maximum error across many widths in practical CodeRM pools.

By Ricardo Fitas
arXiv Machine Learning
Aug 19

Which CS1 Students Will Fail? Identifying Digital Markers from Learning Analytics in Computer Systems and Architecture Using Weighted Academic Momentum and Interaction Logs

The study explores whether combining traditional and digital learning analytics can predict failure in a first‑year CS1 course. Using data from 284 students across four cohorts, the authors identified ten candidate factors and built a logistic regression model that achieved 74.7% accuracy and 0.742 macro F1, with 87% recall for failing students. Weighted academic momentum, basic demographics, and LMS activity emerged as the most predictive features, suggesting that simple digital markers can enable early‑warning systems by week five.

By Lighton Phiri, Mutune Chaibela, Ivy Chisha, David Pungwa, Danny Siabbaba, Bydon Simukoko
Hugging Face Trending Papers
Sep 24

A Risk-Adaptive and Evidence-Constrained Framework for Generative AI Feedback in Programming Education

The paper presents a risk‑adaptive, evidence‑constrained framework that uses learning analytics to provide personalized feedback in introductory programming. By training models on 2,993 failed‑submission states from 215 students, the authors predict persistent failure and generate four tailored feedback conditions for 136 cases. A calibrated risk policy selects interventions for 17.8% of eligible states, capturing 25.2% of persistent failures, and the framework ensures that generated messages contain all required components after evidence gating.