arXiv Machine Learning

Stable and Faithful Explanations for Knowledge Tracing

The paper introduces a validation protocol for knowledge‑tracing models that jointly assesses predictive performance, explanation stability, and faithfulness. Using engineered behavioral features from ASSISTments data, the authors compare an XGBoost model explained with TreeSHAP against four deep‑learning baselines, finding comparable predictive accuracy when information is matched and demonstrating that TreeSHAP rankings are stable and impactful. The study highlights how data preprocessing (e.g., rebuilding the 2009 dataset) can affect both model performance and explanation outcomes.

arXiv AI
4d ago

Learn Now, Use Next, Trust Later: Prequential Test-Time Learning for LLM Agents

The paper introduces StepLearn, a nonparametric framework for prequential test‑time learning in large language model agents. StepLearn separates immediate use of informative transitions from persistent trust, turning each transition into a hypothesis that guides the next step and only reusing it after prospective validation across episodes. Experiments on WebArena‑Lite and ALFWorld show StepLearn improves success rates by 2.2–12.7 percentage points over the strongest baseline, with benefits evident from the first task attempts.

By Tong Zhao, Reed Li, Yuyang Hu, Yutao Zhu, Haijin Liang, Haibo Shi, Yu Lu, Zhicheng Dou
arXiv AI
Jun 30

Deterministic Decisions for High-Stakes AI. A Zero-Egress Pipeline with the Deployability of RAG and the Accuracy of Machine Learning

arXiv:2606. 29280v1 Announce Type: cross Abstract: We identify intervention bias as a previously unquantified failure mode of zero-shot large-language-model (LLM) educational advisory agents: without task-specific training, they recommend action when a hindsight-optimal oracle policy mandates inaction.

By Craig Atkinson
arXiv Machine Learning
Sep 17

No Usable Linear "Capitulation Direction" in Two Small LLMs: A Validation Protocol for Activation-Steering Claims, and a Cross-Family Behavioral Study of Sycophancy Under Pushback

The study examines how two small instruction‑tuned language models, Qwen2.5‑1.5B and Llama‑3.2‑1B, respond to user pushback on TriviaQA. When initially correct, the models flip to a wrong answer in about 42–43% of cases, with the effectiveness of different pushback styles varying by model. Attempts to decode capitulation from the pre‑response residual stream fail under a rigorous validation protocol, revealing overfitting and a measurement hazard that underestimates capitulation by 18–24 percentage points.

By Saad Aamir, Muhammad Awais Bin Adil
arXiv Machine Learning
Aug 19

Study-Strategy Clusters from EdNet Logs Track Engagement, Not Mastery

The study clusters 5,000 EdNet-KT3 learners into eight study‑strategy groups based on early‑session behaviors such as resource use, revision, video watching, and problem practice. These clusters predict later engagement metrics—like continued practice and session completion—but do not reliably forecast later unassisted accuracy or mastery. The findings suggest that behavioral clustering captures learning styles and engagement patterns rather than knowledge gains.

By Qingchuan Lyu, Yingxin Li, Albert Yang