arXiv AI By Liliana Santos-Deonizio, James Malamut, Ram\'on Mart\'inez, Dorottya Demszky

When Youth Enter The Chat: An Epistemic Shift in the Validation of LLM-Based Measures of Student Talk

Read the original on arXiv AI →

The article discusses the growing use of large language models (LLMs) to assess student discourse at scale, noting that current validation methods—such as expert annotations and F1 scores—often ignore the contextual and cultural nuances of student language. It argues that these practices inadequately capture the experiences of racially and linguistically marginalized youth, and proposes re‑contextualizing classroom conversations and involving students as epistemic authorities. A case study with multilingual 8th‑grade math students demonstrates misalignments between student self‑interpretations and LLM outputs, underscoring the need for youth participation in validating LLM‑based measures.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Aug 27

EduDial: Constructing a Large-scale Multi-turn Teacher-Student Dialogue Corpus

EduDial is a large-scale multi-turn teacher‑student dialogue corpus covering 345 core knowledge points and 34,250 dialogue sessions, designed around Bloom’s taxonomy and ten questioning strategies such as situational, ZPD, and metacognitive questioning. The dataset includes differentiated teaching strategies for students at varying cognitive levels to provide targeted guidance. Using EduDial, the authors trained EduDial‑LLM 32B and introduced an 11‑dimensional evaluation framework that measures teaching quality and content quality, showing that most mainstream LLMs struggle with student‑centered teaching while EduDial‑LLM outperforms all baselines across all metrics.

By Shouang Wei, Min Zhang, Xin Lin, Bo Jiang, Zhongxiang Dai, Kun Kuang
arXiv AI
Sep 7

Who Should Grade My Work? Student Perspectives on Transparent AI-Assisted Writing Assessment in Higher Education

The study explores how undergraduate computing students in Saudi Arabia perceive AI‑generated writing feedback when they are explicitly told that ChatGPT, not a human instructor, produced the score and comments. Through qualitative reflections, four themes emerged: students found the feedback useful for surface‑level revisions, recognized AI’s contextual and pedagogical limits, trusted the feedback conditionally—separating its utility from its authority—and reaffirmed the human instructor’s role as the ultimate grading authority. The findings highlight a clear distinction students make between feedback usefulness and evaluative authority, treating them as separate judgments rather than opposing ends of a single approval scale.

By Rayed AlGhamdi