arXiv Machine Learning

Participation-Sensitive Convergence and the Fragment First, Converge Later Pattern in Asynchronous Online Learning: A Topological Analysis Across 22 OULAD Courses

The study analyzes 22 OULAD asynchronous online courses (over 22,000 learners) using Zigzag Persistent Homology to track the number of disconnected behavioral clusters, denoted $eta_0$. It finds that changes in $eta_0$ strongly co‑vary with active learner counts, indicating that $eta_0$ is a participation‑sensitive indicator rather than a cause of dropout. Assessment deadlines trigger fragmentation in 82.6% of cases and the full Fragment First, Converge Later cycle in 60.2%, with long‑term fragmentation dominating 90.9% of courses.

arXiv Machine Learning
Aug 19

Study-Strategy Clusters from EdNet Logs Track Engagement, Not Mastery

The study clusters 5,000 EdNet-KT3 learners into eight study‑strategy groups based on early‑session behaviors such as resource use, revision, video watching, and problem practice. These clusters predict later engagement metrics—like continued practice and session completion—but do not reliably forecast later unassisted accuracy or mastery. The findings suggest that behavioral clustering captures learning styles and engagement patterns rather than knowledge gains.

By Qingchuan Lyu, Yingxin Li, Albert Yang
arXiv Machine Learning
Aug 19

Which CS1 Students Will Fail? Identifying Digital Markers from Learning Analytics in Computer Systems and Architecture Using Weighted Academic Momentum and Interaction Logs

The study explores whether combining traditional and digital learning analytics can predict failure in a first‑year CS1 course. Using data from 284 students across four cohorts, the authors identified ten candidate factors and built a logistic regression model that achieved 74.7% accuracy and 0.742 macro F1, with 87% recall for failing students. Weighted academic momentum, basic demographics, and LMS activity emerged as the most predictive features, suggesting that simple digital markers can enable early‑warning systems by week five.

By Lighton Phiri, Mutune Chaibela, Ivy Chisha, David Pungwa, Danny Siabbaba, Bydon Simukoko
arXiv Computation and Language
Sep 22

Time-Incremental Continued Pretraining of LLMs: Knowledge Updates Without Catastrophic Forgetting

The paper investigates time‑incremental continued pretraining (CPT) of large language models using web‑scale data that overlaps across snapshots. Across six open‑weight models, CPT improves factual recall without catastrophic forgetting, while the cost is negligible and data quality outweighs quantity. The study identifies optimal learning rates, shows LoRA can match full CPT, and demonstrates that CPT gains transfer to fine‑tuned models.

By F{\i}rat \"Oncel, Salman Hussain Ali, Mirco Ravanelli, Cem Subakan, \c{C}a\u{g}atay Y{\i}ld{\i}z