EduDial is a large-scale multi-turn teacher‑student dialogue corpus covering 345 core knowledge points and 34,250 dialogue sessions, designed around Bloom’s taxonomy and ten questioning strategies such as situational, ZPD, and metacognitive questioning. The dataset includes differentiated teaching strategies for students at varying cognitive levels to provide targeted guidance. Using EduDial, the authors trained EduDial‑LLM 32B and introduced an 11‑dimensional evaluation framework that measures teaching quality and content quality, showing that most mainstream LLMs struggle with student‑centered teaching while EduDial‑LLM outperforms all baselines across all metrics.
By Shouang Wei, Min Zhang, Xin Lin, Bo Jiang, Zhongxiang Dai, Kun Kuang
arXiv:2604. 21137v3 Announce Type: replace-cross Abstract: Analyzing the reasoning patterns of students in science classrooms is critical for understanding knowledge construction mechanism and improving instructional practice to maximize cognitive engagement, yet manual coding of classroom discourse at scale remains prohibitively labor-intensive.
By Jiho Noh, Mukhesh Raghava Katragadda, Raymond Carl, Soon Lee
arXiv:2609.16993v1 Announce Type: cross
Abstract: Large Language Models are now common in student assessment, but we know little about how student demographics affect their use. Sometimes, considerin...
By Donya Rooein, Luca Benedetto, Dirk Hovy
PROOF is a benchmark that profiles the reliability of object-level facts in instruction-tuned language models by converting a frozen Wikidata snapshot into 18,486 English multiple-choice questions grounded in 11,779 semantic facts across 101 classes, 392 properties, and 14 domains. Each question includes an explicit "I don't know" option, a "No correct option" control, and nine controlled formulations, with 1,849 questions designed as no-correct-option traps. The study evaluates 18 open-weight model deployments on 166,374 prompts, revealing wide variability in factual accuracy, sensitivity to wording changes, and the impact of decoder perturbations.
By Andrei Chetvergov, Mikhail Solovev, Timofei Sivoraksha, Stepan Ukolov, Valeriia Kuschenko, Alexander Evseev, Sergey Bolovtsov
arXiv:2602. 23971v4 Announce Type: replace-cross Abstract: Sycophancy, the tendency of large language models to favour user-affirming responses over critical engagement, has been identified as an alignment failure, particularly in high-stakes advisory and social contexts.
By Magda Dubois, Cozmin Ududec, Christopher Summerfield, Lennart Luettgau
arXiv:2609.08016v1 Announce Type: new
Abstract: Multi-agent debate, in which several LLMs exchange arguments before answering, is widely assumed to improve answer quality by surfacing genuine disagre...
By Chen Qian
arXiv:2603.17839v4 Announce Type: replace-cross
Abstract: Verbal confidence -- prompting LLMs to state their confidence as a number or category -- is widely used to extract uncertainty estimates from...
By Dharshan Kumaran, Arthur Conmy, Federico Barbero, Simon Osindero, Viorica Patraucean, Petar Veli\v{c}kovi\'c
arXiv:2602. 09907v3 Announce Type: replace-cross Abstract: College students increasingly use AI chatbots to support academic reading, yet we lack granular understanding of how these interactions shape their reading experience and cognitive engagement.
By Yue Fu, Joel Wester, Niels Van Berkel, Alexis Hiniker
The paper investigates how the way users phrase advice‑seeking requests—termed articulation—creates stable, measurable patterns distinct from the topics of the requests. By analyzing 16,447 prompts from public chat corpora, the authors identify a small set of latent articulation factors that consistently appear across datasets and splits. One key finding is a long‑form, information‑poor style that leads language models to give shorter, vaguer answers without seeking clarification, a pattern that persists across topics and prompt lengths.
By Juneha Baek, Suhyeon Lee, Donghyuk Shin
arXiv:2605. 21629v2 Announce Type: replace-cross Abstract: How much have students' ordinary learning processes shifted in response to generative AI, and how does that affect their durable learning outcomes?
By Sina Rismanchian, Hasan Uzun, Jeffrey Matayoshi, Eric Cosyn, Eyad Kurd-Misto
The study analyzes 20,462 student turns from 1,260 sessions with a guided LLM chemistry tutor, identifying 6,630 impasse turns categorized as conceptual errors, expressed uncertainty, or help‑seeking. Three tutoring conditions—baseline, no‑direct‑answer, and guided—were simulated, revealing that the baseline tutor often gave direct answers, the no‑direct‑answer tutor always asked follow‑up questions, and the guided tutor varied its responses based on context. Impasse trajectories showed that each additional impasse turn reduced the likelihood of recovery, while addressing errors became increasingly beneficial compared to repeated scripted questioning.
By Bakhtawar Ahtisham, Kirk Vanacore, Alessandra Napoli, Josh Arens, Ksenia Ionova, Clayton Cohn, Shima Salehi, Rene Kizilcec
The paper investigates token‑level certainty as a proxy for correctness in large language models. It finds that certainty better predicts whether a model will answer a question correctly than it does whether a specific response is correct, and that certainty varies by token type and position. The authors show that using certainty early in generation to allocate responses and later to weight votes improves accuracy while dramatically cutting token cost.
By Yunfan Zhou, Ye Zhu, Zhihai Wang, Jianguo Yao, Haibing Guan, Xijun Li