arXiv AI

Characterizing Questioning Patterns and Student Engagement Through Contextual Analysis of Real-Time Classroom Interactions

The study analyzes 604 real‑time classroom poll questions from 47 live sessions, aligning each poll with lecture transcripts and attendance data. It finds that 89% of poll answers can be located in the lecture context, revealing that many polls serve attention‑checking functions only visible when contextualized. The majority of questions are lower‑order and fit into seven instructional functions, with student engagement high overall but uneven, and students often misjudge their own correctness.

arXiv Computation and Language
Aug 27

EduDial: Constructing a Large-scale Multi-turn Teacher-Student Dialogue Corpus

EduDial is a large-scale multi-turn teacher‑student dialogue corpus covering 345 core knowledge points and 34,250 dialogue sessions, designed around Bloom’s taxonomy and ten questioning strategies such as situational, ZPD, and metacognitive questioning. The dataset includes differentiated teaching strategies for students at varying cognitive levels to provide targeted guidance. Using EduDial, the authors trained EduDial‑LLM 32B and introduced an 11‑dimensional evaluation framework that measures teaching quality and content quality, showing that most mainstream LLMs struggle with student‑centered teaching while EduDial‑LLM outperforms all baselines across all metrics.

By Shouang Wei, Min Zhang, Xin Lin, Bo Jiang, Zhongxiang Dai, Kun Kuang
arXiv AI
Aug 18

Enhancing Science Classroom Discourse Analysis through Joint Multi-Task Learning for Reasoning-Component Classification

arXiv:2604. 21137v3 Announce Type: replace-cross Abstract: Analyzing the reasoning patterns of students in science classrooms is critical for understanding knowledge construction mechanism and improving instructional practice to maximize cognitive engagement, yet manual coding of classroom discourse at scale remains prohibitively labor-intensive.

By Jiho Noh, Mukhesh Raghava Katragadda, Raymond Carl, Soon Lee
arXiv AI
Sep 25

PROOF: Profiling Reliability of Object-Level Facts in Large Language Models

PROOF is a benchmark that profiles the reliability of object-level facts in instruction-tuned language models by converting a frozen Wikidata snapshot into 18,486 English multiple-choice questions grounded in 11,779 semantic facts across 101 classes, 392 properties, and 14 domains. Each question includes an explicit "I don't know" option, a "No correct option" control, and nine controlled formulations, with 1,849 questions designed as no-correct-option traps. The study evaluates 18 open-weight model deployments on 166,374 prompts, revealing wide variability in factual accuracy, sensitivity to wording changes, and the impact of decoder perturbations.

By Andrei Chetvergov, Mikhail Solovev, Timofei Sivoraksha, Stepan Ukolov, Valeriia Kuschenko, Alexander Evseev, Sergey Bolovtsov
arXiv AI
Jul 31

Ask don't tell: Reducing sycophancy in large language models

arXiv:2602. 23971v4 Announce Type: replace-cross Abstract: Sycophancy, the tendency of large language models to favour user-affirming responses over critical engagement, has been identified as an alignment failure, particularly in high-stakes advisory and social contexts.

By Magda Dubois, Cozmin Ududec, Christopher Summerfield, Lennart Luettgau
arXiv AI
Sep 21

How do LLMs Compute Verbal Confidence

arXiv:2603.17839v4 Announce Type: replace-cross Abstract: Verbal confidence -- prompting LLMs to state their confidence as a number or category -- is widely used to extract uncertainty estimates from...

By Dharshan Kumaran, Arthur Conmy, Federico Barbero, Simon Osindero, Viorica Patraucean, Petar Veli\v{c}kovi\'c
arXiv Computation and Language
Sep 1

How You Ask Shapes What You Get: A Theory-Seeded Measurement of Articulation in Advice-Seeking LLM Conversations

The paper investigates how the way users phrase advice‑seeking requests—termed articulation—creates stable, measurable patterns distinct from the topics of the requests. By analyzing 16,447 prompts from public chat corpora, the authors identify a small set of latent articulation factors that consistently appear across datasets and splits. One key finding is a long‑form, information‑poor style that leads language models to give shorter, vaguer answers without seeking clarification, a pattern that persists across topics and prompt lengths.

By Juneha Baek, Suhyeon Lee, Donghyuk Shin
arXiv AI
3d ago

Examining Variation in How Guided AI Tutors Resolve Student Impasses

The study analyzes 20,462 student turns from 1,260 sessions with a guided LLM chemistry tutor, identifying 6,630 impasse turns categorized as conceptual errors, expressed uncertainty, or help‑seeking. Three tutoring conditions—baseline, no‑direct‑answer, and guided—were simulated, revealing that the baseline tutor often gave direct answers, the no‑direct‑answer tutor always asked follow‑up questions, and the guided tutor varied its responses based on context. Impasse trajectories showed that each additional impasse turn reduced the likelihood of recovery, while addressing errors became increasingly beneficial compared to repeated scripted questioning.

By Bakhtawar Ahtisham, Kirk Vanacore, Alessandra Napoli, Josh Arens, Ksenia Ionova, Clayton Cohn, Shima Salehi, Rene Kizilcec
arXiv AI
2d ago

Certainty Is Not Just Correctness: Rethinking Token-Level Certainty in LLM Reasoning

The paper investigates token‑level certainty as a proxy for correctness in large language models. It finds that certainty better predicts whether a model will answer a question correctly than it does whether a specific response is correct, and that certainty varies by token type and position. The authors show that using certainty early in generation to allocate responses and later to weight votes improves accuracy while dramatically cutting token cost.

By Yunfan Zhou, Ye Zhu, Zhihai Wang, Jianguo Yao, Haibing Guan, Xijun Li