arXiv Computation and Language

Modeling Student Sensemaking with LLMs and Knowledge-Graph-Guided Inference

The study explores whether instruction-tuned large language models (LLMs) can analyze collaborative student sensemaking without task-specific training, and whether adding structured knowledge-state information enhances this analysis. Two mid-size LLMs were evaluated on 23 expert-labeled episodes under various prompting conditions, showing that reasoning-enabled prompts better detect unsuccessful sensemaking and that knowledge-state diagnostics improve agreement with experts. No single configuration outperformed others across all sensemaking dimensions, highlighting the task’s multidimensional nature.

arXiv Computation and Language
Aug 27

EduDial: Constructing a Large-scale Multi-turn Teacher-Student Dialogue Corpus

EduDial is a large-scale multi-turn teacher‑student dialogue corpus covering 345 core knowledge points and 34,250 dialogue sessions, designed around Bloom’s taxonomy and ten questioning strategies such as situational, ZPD, and metacognitive questioning. The dataset includes differentiated teaching strategies for students at varying cognitive levels to provide targeted guidance. Using EduDial, the authors trained EduDial‑LLM 32B and introduced an 11‑dimensional evaluation framework that measures teaching quality and content quality, showing that most mainstream LLMs struggle with student‑centered teaching while EduDial‑LLM outperforms all baselines across all metrics.

By Shouang Wei, Min Zhang, Xin Lin, Bo Jiang, Zhongxiang Dai, Kun Kuang
arXiv Computation and Language
4d ago

Layer-Informed Fine-Tuning via Three-Stage Functional Segmentation of LLMs

The paper proposes Layer-Informed Fine-Tuning (LIFT), a method that identifies and updates only the most functionally critical layers of large language models (LLMs) using a bottleneck identification mechanism based on sensitivity analysis. By focusing on layers that handle conceptualization, reasoning, and textualization, LIFT aims to accelerate training and enhance performance on reasoning tasks. Experiments demonstrate that this selective fine-tuning approach both speeds up the training process and yields significant performance gains.

By Junning Shao, Siwei Wang, Zhixuan Fang
arXiv AI
Jul 8

LongCrafter: Towards Diverse Long-Context Understanding via Evidence-Graph-Guided Instruction Synthesis

arXiv:2607. 06160v1 Announce Type: cross Abstract: Synthesizing long-context supervised fine-tuning (SFT) data is a scalable way to enhance the long-context understanding of large language models (LLMs), yet existing approaches share three limitations: narrow task coverage, insufficient instruction difficulty, and a lack of faithfulness supervision.

By Chenhao Yuan, Yinhao Xu, Shuwen Xu, Xizhi Yang, Jiaxiang Liu, Chenxi Zhou, Shaoping Huang, Haolin Ren, Pengfei Cao, Jun Zhao, Kang Liu
arXiv AI
Jul 24

AI Assistants Overassist

arXiv:2607. 21306v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as tutors and thought partners, helping users reason through problems.

By Verona Teo, Raghav Jain, Tobias Gerstenberg, Max Kleiman-Weiner
arXiv AI
Sep 24

Experts Rise Where LLMs Disagree: Using Cross-Model Disagreement to Target Expert Effort in LLM Codebook Revision for Large-Scale Annotation

The paper proposes using large language models (LLMs) to identify disagreements among models as a way to focus expert effort on revising codebooks for large‑scale text annotation. Three expert feedback methods are evaluated: editing LLM‑generated revisions (Codebook Verifying), answering questions about disagreements (Question Answering), and labeling disagreement cases with rationales (Rationale Labeling). Experiments on tutoring‑session transcripts show that Rationale Labeling achieves the highest LLM‑labeling accuracy (64.9%) compared to the expert‑revised codebook (57.8%), with Question Answering also outperforming the baseline (60.5%).

By Zeyu He, Zhuqian Zhou, Kirk Vanacore, Rene F. Kizilcec, Ting-Hao 'Kenneth' Huang