The study explores whether combining traditional and digital learning analytics can predict failure in a first‑year CS1 course. Using data from 284 students across four cohorts, the authors identified ten candidate factors and built a logistic regression model that achieved 74.7% accuracy and 0.742 macro F1, with 87% recall for failing students. Weighted academic momentum, basic demographics, and LMS activity emerged as the most predictive features, suggesting that simple digital markers can enable early‑warning systems by week five.
By Lighton Phiri, Mutune Chaibela, Ivy Chisha, David Pungwa, Danny Siabbaba, Bydon Simukoko
arXiv:2604. 08874v3 Announce Type: replace-cross Abstract: This study proposes a temporal modeling framework with a counterfactual policy-simulation layer for student dropout in higher education, using LMS engagement data and administrative withdrawal records.
By Rafael da Silva, Jeff Eicher, Gregory Longo
arXiv:2608. 04408v1 Announce Type: cross Abstract: On-policy distillation (OPD) supervises student-visited trajectories, yet divergence-based rules cannot determine whether an erroneous prefix remains correctable.
By De Jiang, Zhengyang Zhang, Kehong Yuan, Shaohua Ma
arXiv:2608.29517v1 Announce Type: new
Abstract: Large language models (LLMs) are increasingly used as essay graders in learning analytics, evaluated almost exclusively with agreement statistics. Educ...
By Veerendra Kumar Sunkavalli
The study evaluates whether large language models (LLMs) with in‑context learning can better identify institution‑specific protected health information (PHI) in electronic health records than existing de‑identification systems. Using 100 pediatric oncology notes from Texas Children’s Hospital, eight LLMs were compared to two purpose‑built systems and pattern‑based baselines under three progressively specific prompts. The best LLM achieved an F1 score of 0.918, recovering 79% of previously missed PHI categories and reaching a recall of 0.981 after iterative prompt refinement, demonstrating that calibrated single‑pass prompting can close the institutional PHI gap while balancing precision and recall.
By Daniel Palacios, Matthew Brady Neeley, Angel Adetomike Otto, Shalini Dhamodharan, John P. Woodhouse, Chi-fan Lin, Mark Zobeck, Zhandong Liu, Hyun-Hwan Jeong
The paper reports on a series of experiments examining how different forms of directives—such as record pointers, criteria, or combinations—affect an agent’s choice of archived source records when it inherits six one-line memories. Across twelve registered studies involving 14,760 attempts on a single instrument lineage, the authors measured the impact of various directive formats on six direct-provider models, nine OpenRouter-served models, and several Claude and Opus 5 models, noting differences in performance metrics and replication outcomes. The results are purely descriptive, detailing the effects of exact edits on fixed panels with registered intervals and no claim of underlying mechanisms.
By Kazuki Nakayashiki
arXiv:2605. 21629v2 Announce Type: replace-cross Abstract: How much have students' ordinary learning processes shifted in response to generative AI, and how does that affect their durable learning outcomes?
By Sina Rismanchian, Hasan Uzun, Jeffrey Matayoshi, Eric Cosyn, Eyad Kurd-Misto
arXiv:2606. 19469v1 Announce Type: new Abstract: Undergraduate computer science is governed by international curricular guidelines revised about once a decade, yet programs lack a reliable, reproducible way to measure how completely they cover the current guidelines and how that coverage shifts when the guidelines are restructured.
By Sherzod Turaev, Mary John, Saja Aldabet, Mamoun Awad, Nazar Zaki, Khaled Shuaib
arXiv:2608. 14509v1 Announce Type: new Abstract: Systems that ask a language model to reach a conclusion from many sources usually concatenate them into one prompt.
By Zhelun Wu
arXiv:2609.38021v1 Announce Type: cross
Abstract: We evaluate an auditable long-term memory system on LongMemEval-S. Its retrieval chain uses hybrid candidate retrieval, cross-encoder reranking, cove...
By Christopher J. Chanhnourack
The study evaluates large language model (LLM) graders on two computer‑science exams, testing 171 configurations of closed‑ and open‑weights models. While the best LLM configuration achieved a mean absolute error of 1.64/35—better than the 2.61/35 error between two human graders—its performance was highly sensitive to the prompt. A short "strict grader" preamble caused most open‑weight models to exceed acceptable error thresholds or stop grading entirely, whereas fine‑tuning with a single LoRA adapter restored parity with human graders and reduced sensitivity to harsh prompts.
By Ali Habibullah, Yazan Alshoibi, Mohammad Alshiekh, Salman Khan, Naeemullah Khan
arXiv:2609.24194v1 Announce Type: new
Abstract: Evaluation scores used around LLM systems -- including reward models, rerankers, and LLM judges -- can track surface form instead of the quality they c...
By Daein Weon, Dongho Kang