The study evaluates how well pre‑trained models can classify the Bloom level of AI‑generated educational questions, a task that is crucial for ensuring pedagogical quality. Traditional machine‑learning models perform poorly on out‑of‑distribution data, whereas transformer and large‑language models achieve higher accuracy, especially after feature‑engineering techniques such as text splicing and appending learning objectives. Retraining the models yields the most significant performance gains across all datasets.
By Michael Lawrence Castanares, Princess Ventures, Allan Tan
arXiv:2603. 02830v2 Announce Type: replace-cross Abstract: Predicting future student responses to questions is particularly valuable for educational learning platforms where it enables effective interventions.
By Prarthana Bhattacharyya, Joshua Mitton, Ralph Abboud, Simon Woodhead
arXiv:2606. 18257v1 Announce Type: cross Abstract: While LLMs show promise in automating educational content creation, their ability to generate questions that stimulate higher-order thinking remains understudied.
By Xiaolong Wang, Zhe Zhao, Song Lai, Chaoli Zhang, Zijie Geng, Yu Tong, Ye Wei, Qingsong Wen
EduDial is a large-scale multi-turn teacher‑student dialogue corpus covering 345 core knowledge points and 34,250 dialogue sessions, designed around Bloom’s taxonomy and ten questioning strategies such as situational, ZPD, and metacognitive questioning. The dataset includes differentiated teaching strategies for students at varying cognitive levels to provide targeted guidance. Using EduDial, the authors trained EduDial‑LLM 32B and introduced an 11‑dimensional evaluation framework that measures teaching quality and content quality, showing that most mainstream LLMs struggle with student‑centered teaching while EduDial‑LLM outperforms all baselines across all metrics.
By Shouang Wei, Min Zhang, Xin Lin, Bo Jiang, Zhongxiang Dai, Kun Kuang
arXiv:2606. 12767v1 Announce Type: new Abstract: Evaluating procedural reasoning in AI-supported learning systems requires question-answer datasets that are both learner-like and grounded in the instructional knowledge the system is expected to use.
By Sarah Elshabrawy, Rahul K. Dass, Ashok K. Goel
The paper introduces a prompt‑engineering framework that personalizes large language model (LLM) teaching assistants across disciplines by tailoring responses to six learner‑specific dimensions, creating 96 distinct learner profiles. It also analyzes student queries through Bloom’s Taxonomy to gauge cognitive complexity, encoding both learner attributes and cognitive assessments into structured prompts that condition the LLM without retraining. Experiments using NLP metrics and a small human study demonstrate that this approach yields perceptible differences in response style and structure, with statistical evidence linking specific learner attributes to measurable changes.
By Saptarshi Basu, Sandeep Kakar, Ashok Goel