arXiv Computation and Language By Oliver G. B. Garrod, Robin A. A. Ince, Meng Liu, Mohamed Huti, Moritz Boos, Amy Waldock, Dominic Andrews, Paul Atherton

Edu-QuRating: Multi-Dimensional Educational Data Curation with Distilled Pairwise Judgements

Read the original on arXiv Computation and Language →

Edu-QuRating is a pipeline that scores and curates educational data across multiple dimensions—accuracy, engagement, structure, and audience appropriateness—using an LLM judge to label document pairs and distill these preferences into reusable Edu-QuRaters. The best Edu-QuRater achieves 91.7% accuracy against held‑out GPT‑4.1‑mini judgments and is applied to filter 322.25 M FineWeb‑Edu‑Fortified documents, improving small‑model pre‑training performance on nine benchmarks. Additionally, Edu-QuRater scores serve as reward signals in GRPO post‑training, yielding responses that are preferred for pedagogical quality and instruction following over the Qwen3‑4B base model.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Sep 24

Evaluation of pre-trained models for pedagogical assessment of novel AI-assisted educational questions

The study evaluates how well pre‑trained models can classify the Bloom level of AI‑generated educational questions, a task that is crucial for ensuring pedagogical quality. Traditional machine‑learning models perform poorly on out‑of‑distribution data, whereas transformer and large‑language models achieve higher accuracy, especially after feature‑engineering techniques such as text splicing and appending learning objectives. Retraining the models yields the most significant performance gains across all datasets.

By Michael Lawrence Castanares, Princess Ventures, Allan Tan
arXiv AI
Aug 25

AI University: An LLM-Powered Learning Assistant for Engineering---A Finite Element Method Case Study

AI University (AI‑U) is a flexible framework that uses a fine‑tuned large language model (LLM) combined with retrieval‑augmented generation (RAG) and a reasoning synthesis model to produce style‑aligned responses from lecture videos, notes, and textbooks. In a graduate‑level finite‑element‑method (FEM) course, the authors created a pipeline to generate course‑grounded training data, fine‑tuned an open‑source LLM with Low‑Rank Adaptation (LoRA), and applied RAG‑based synthesis. Evaluation through cosine similarity, LLM‑based assessment, expert review, and user studies showed that the expert model outperformed the base model in alignment with course materials, with 86 % of test cases scoring higher and human users preferring the expert model roughly twice as often. whyItMatters":"The study demonstrates a practical method for building course‑specific learning assistants that improve alignment with instructional content, offering a template that can be extended across STEM fields."

By Mostafa Faghih Shojaei, Rahul Gulati, Benjamin A. Jasperson, Shangshang Wang, Simone Cimolato, Manas Vardhan, Dangli Cao, Willie Neiswanger, Krishna Garikipati
arXiv Machine Learning
Sep 11

E-CONAN (Entailment, CONtradition And Neutral) Benchmarks: Arabic Textual Entailment and Natural Inference Datasets

E-CONAN introduces Arabic textual entailment and natural inference benchmarks comprising two datasets: E-CONAN-2 (2-way RTE) and E-CONAN-3 (3-way NLI). The datasets are built from automatically-translated pairs, human-validated machine translations, hand-crafted pairs from Arabic teaching books, and rumor-containing news headlines. The authors evaluated nine multilingual pretrained models and five large language models on these benchmarks, demonstrating that E-CONAN offers a more diverse and robust assessment than existing datasets like XNLI and ArNLI.

By Khloud AL Jallad, Nada Ghneim, Ghaida Rebdawi