arXiv AI

A Multi-level Analysis of Factors Associated with Student Performance: A Machine Learning Approach to the SAEB Microdata

arXiv:2510. 22266v3 Announce Type: replace-cross Abstract: Identifying the factors that influence student performance in basic education is a central challenge for formulating effective public policies in Brazil.

arXiv Machine Learning
Sep 23

A Hybrid AI Framework for Academic Advising: Integrating Ensemble-Based Grade Prediction and a Rule-Based Expert System

The paper presents a hybrid AI framework for academic advising that combines an ensemble-based grade prediction model with a rule‑based expert system. Students from the University of Birjand were clustered using Gaussian Mixture Models, and a Stacking Ensemble of Random Forest, Gradient Boosting, and MLP was trained per cluster, achieving an aggregated RMSE of 2.35. The expert system then uses these predictions alongside institutional regulations to give real‑time feedback such as GPA forecasts, probation warnings, and course recommendations.

By Hamid Saadatfar, Rohollah Hedayati-Nasab, AmirHossein Eshghi, Arash Hajihashemi
arXiv Machine Learning
Jul 30

Archetypes or ability? Clustering for modelling student mathematical competence

arXiv:2607. 26063v1 Announce Type: cross Abstract: Personalised learning systems often assume that mathematical ability is combined of discrete abilities, acquired sequentially and dependent upon first acquiring foundational abilities, and students often report different strengths.

By Benjamin Mawdsley, Tom Quilter, Richard Turner, Sarah Jackson, Paul Edwards
arXiv Machine Learning
Aug 19

Predicting Male Domestic Violence Using Explainable Ensemble Learning and Exploratory Data Analysis

The paper presents a data‑driven study of male domestic violence (MDV) in Bangladesh, using exploratory data analysis to uncover patterns such as verbal abuse prevalence and the influence of financial dependency. It evaluates 10 traditional ML models, 3 deep learning models, and 2 ensemble models, ultimately proposing a stacking ensemble with ANN and CatBoost base classifiers and Logistic Regression meta‑model that achieves 95% accuracy and 99.29% AUC. Explainable AI techniques (SHAP, LIME) and statistical validation confirm the model’s superior performance and highlight key features driving predictions.

By Md Abrar Jahin, Saleh Akram Naife, Fatema Tuj Johora Lima, M. F. Mridha, Md. Jakir Hossen
arXiv AI
Jul 3

AIriskEval-edu: New Dataset for Risk Assessment in AI-mediated K-12 Educational Explanations

arXiv:2607. 01934v1 Announce Type: cross Abstract: This work introduces AIriskEval-edu-db2, a new dataset designed to train and evaluate auditors based on LLMs for an explainable pedagogical risk assessment in instructional content for grades K-12.

By Javier Irigoyen, Roberto Daza, Francisco Jurado, Julian Fierrez, Ruben Tolosana, Alvaro Ortigosa, Enrique Blas, Aythami Morales
Hugging Face Trending Papers
Jul 2

AIriskEval-edu: New Dataset for Risk Assessment in AI-mediated K-12 Educational Explanations

This work introduces AIriskEval-edu-db2, a new dataset designed to train and evaluate auditors based on LLMs for an explainable pedagogical risk assessment in instructional content for grades K-12. The dataset comprises 1,639 explanations from 170 curated ScienceQA questions, covering science, language arts, and social sciences.

arXiv Machine Learning
Sep 25

Limited Structural Reliability in Public Educational Prediction Benchmarks: A Four-Dimension Audit of Seven Datasets

The study audited seven public educational prediction datasets using four pre‑modeling reliability checks—baseline gap, split instability, null separation, and metadata adequacy under group‑aware holdout. Only three datasets passed all checks; the others failed either group‑aware generalization tests or lacked necessary provenance metadata. The audit revealed that cross‑group fragility, rather than weak iid performance, was the dominant failure mode, and that increasing model complexity did not resolve these structural issues.

By Yan Ma, Lizhuo Zhang
arXiv Machine Learning
Aug 19

Which CS1 Students Will Fail? Identifying Digital Markers from Learning Analytics in Computer Systems and Architecture Using Weighted Academic Momentum and Interaction Logs

The study explores whether combining traditional and digital learning analytics can predict failure in a first‑year CS1 course. Using data from 284 students across four cohorts, the authors identified ten candidate factors and built a logistic regression model that achieved 74.7% accuracy and 0.742 macro F1, with 87% recall for failing students. Weighted academic momentum, basic demographics, and LMS activity emerged as the most predictive features, suggesting that simple digital markers can enable early‑warning systems by week five.

By Lighton Phiri, Mutune Chaibela, Ivy Chisha, David Pungwa, Danny Siabbaba, Bydon Simukoko