arXiv Machine Learning

OBER+: Continuity-Aware Reporting and Traceable Continuous Improvement in Outcome-Based Education

OBER+ extends an institutional attainment platform to bridge the gap between measured learning outcome shortfalls and corrective actions. It aggregates attainment across course deliveries, flags shortfalls, grades them against regulator cutoffs, records decisions linked to evidence‑annotated practices, logs changes, and quantifies subsequent improvements. The system also ensures outcomes are compared only when unchanged, preventing misleading comparisons across redefined outcomes, and has identified real defects in institutional data through rule‑based analysis.

arXiv Machine Learning
Aug 19

Which CS1 Students Will Fail? Identifying Digital Markers from Learning Analytics in Computer Systems and Architecture Using Weighted Academic Momentum and Interaction Logs

The study explores whether combining traditional and digital learning analytics can predict failure in a first‑year CS1 course. Using data from 284 students across four cohorts, the authors identified ten candidate factors and built a logistic regression model that achieved 74.7% accuracy and 0.742 macro F1, with 87% recall for failing students. Weighted academic momentum, basic demographics, and LMS activity emerged as the most predictive features, suggesting that simple digital markers can enable early‑warning systems by week five.

By Lighton Phiri, Mutune Chaibela, Ivy Chisha, David Pungwa, Danny Siabbaba, Bydon Simukoko
arXiv AI
Aug 19

Institution-Specific LLM Prompting Recovers PHI That De-identification Systems and Their Gold Standards Both Miss

The study evaluates whether large language models (LLMs) with in‑context learning can better identify institution‑specific protected health information (PHI) in electronic health records than existing de‑identification systems. Using 100 pediatric oncology notes from Texas Children’s Hospital, eight LLMs were compared to two purpose‑built systems and pattern‑based baselines under three progressively specific prompts. The best LLM achieved an F1 score of 0.918, recovering 79% of previously missed PHI categories and reaching a recall of 0.981 after iterative prompt refinement, demonstrating that calibrated single‑pass prompting can close the institutional PHI gap while balancing precision and recall.

By Daniel Palacios, Matthew Brady Neeley, Angel Adetomike Otto, Shalini Dhamodharan, John P. Woodhouse, Chi-fan Lin, Mark Zobeck, Zhandong Liu, Hyun-Hwan Jeong
arXiv AI
Sep 4

Plan Pointers and Record-Directive Form in Budgeted Verification of Inherited Agent Memory

The paper reports on a series of experiments examining how different forms of directives—such as record pointers, criteria, or combinations—affect an agent’s choice of archived source records when it inherits six one-line memories. Across twelve registered studies involving 14,760 attempts on a single instrument lineage, the authors measured the impact of various directive formats on six direct-provider models, nine OpenRouter-served models, and several Claude and Opus 5 models, noting differences in performance metrics and replication outcomes. The results are purely descriptive, detailing the effects of exact edits on fixed panels with registered intervals and no claim of underlying mechanisms.

By Kazuki Nakayashiki
arXiv AI
Jun 19

Measuring Curriculum Alignment across Topical Coverage, Competency, and Cognitive Depth: A Longitudinal Framework Applied to CS2013 and CS2023

arXiv:2606. 19469v1 Announce Type: new Abstract: Undergraduate computer science is governed by international curricular guidelines revised about once a decade, yet programs lack a reliable, reproducible way to measure how completely they cover the current guidelines and how that coverage shifts when the guidelines are restructured.

By Sherzod Turaev, Mary John, Saja Aldabet, Mamoun Awad, Nazar Zaki, Khaled Shuaib
arXiv AI
Sep 25

Where LLM Graders Succeed and Break: Evidence from Two Computer-Science Exams

The study evaluates large language model (LLM) graders on two computer‑science exams, testing 171 configurations of closed‑ and open‑weights models. While the best LLM configuration achieved a mean absolute error of 1.64/35—better than the 2.61/35 error between two human graders—its performance was highly sensitive to the prompt. A short "strict grader" preamble caused most open‑weight models to exceed acceptable error thresholds or stop grading entirely, whereas fine‑tuning with a single LoRA adapter restored parity with human graders and reduced sensitivity to harsh prompts.

By Ali Habibullah, Yazan Alshoibi, Mohammad Alshiekh, Salman Khan, Naeemullah Khan