arXiv Computation and Language

Edu-QuRating: Multi-Dimensional Educational Data Curation with Distilled Pairwise Judgements

Edu-QuRating is a pipeline that scores and curates educational data across multiple dimensions—accuracy, engagement, structure, and audience appropriateness—using an LLM judge to label document pairs and distill these preferences into reusable Edu-QuRaters. The best Edu-QuRater achieves 91.7% accuracy against held‑out GPT‑4.1‑mini judgments and is applied to filter 322.25 M FineWeb‑Edu‑Fortified documents, improving small‑model pre‑training performance on nine benchmarks. Additionally, Edu-QuRater scores serve as reward signals in GRPO post‑training, yielding responses that are preferred for pedagogical quality and instruction following over the Qwen3‑4B base model.

arXiv AI
Sep 24

Evaluation of pre-trained models for pedagogical assessment of novel AI-assisted educational questions

The study evaluates how well pre‑trained models can classify the Bloom level of AI‑generated educational questions, a task that is crucial for ensuring pedagogical quality. Traditional machine‑learning models perform poorly on out‑of‑distribution data, whereas transformer and large‑language models achieve higher accuracy, especially after feature‑engineering techniques such as text splicing and appending learning objectives. Retraining the models yields the most significant performance gains across all datasets.

By Michael Lawrence Castanares, Princess Ventures, Allan Tan
arXiv AI
Aug 25

AI University: An LLM-Powered Learning Assistant for Engineering---A Finite Element Method Case Study

AI University (AI‑U) is a flexible framework that uses a fine‑tuned large language model (LLM) combined with retrieval‑augmented generation (RAG) and a reasoning synthesis model to produce style‑aligned responses from lecture videos, notes, and textbooks. In a graduate‑level finite‑element‑method (FEM) course, the authors created a pipeline to generate course‑grounded training data, fine‑tuned an open‑source LLM with Low‑Rank Adaptation (LoRA), and applied RAG‑based synthesis. Evaluation through cosine similarity, LLM‑based assessment, expert review, and user studies showed that the expert model outperformed the base model in alignment with course materials, with 86 % of test cases scoring higher and human users preferring the expert model roughly twice as often. whyItMatters":"The study demonstrates a practical method for building course‑specific learning assistants that improve alignment with instructional content, offering a template that can be extended across STEM fields."

By Mostafa Faghih Shojaei, Rahul Gulati, Benjamin A. Jasperson, Shangshang Wang, Simone Cimolato, Manas Vardhan, Dangli Cao, Willie Neiswanger, Krishna Garikipati
arXiv Machine Learning
Sep 11

E-CONAN (Entailment, CONtradition And Neutral) Benchmarks: Arabic Textual Entailment and Natural Inference Datasets

E-CONAN introduces Arabic textual entailment and natural inference benchmarks comprising two datasets: E-CONAN-2 (2-way RTE) and E-CONAN-3 (3-way NLI). The datasets are built from automatically-translated pairs, human-validated machine translations, hand-crafted pairs from Arabic teaching books, and rumor-containing news headlines. The authors evaluated nine multilingual pretrained models and five large language models on these benchmarks, demonstrating that E-CONAN offers a more diverse and robust assessment than existing datasets like XNLI and ArNLI.

By Khloud AL Jallad, Nada Ghneim, Ghaida Rebdawi
arXiv Computation and Language
Sep 16

Style-Debiased DPO: Updating LLM Knowledge with Factuality-Aware Synthetic Preference Data

The paper introduces Style‑Debiased DPO (SD‑DPO), a method that refines large language models’ ability to retrieve stored knowledge by using preference optimization that corrects for style differences while preserving factual accuracy. SD‑DPO evaluates on the EntiGraph storing‑side framework and outperforms baseline CPT on the QuALITY reading‑comprehension benchmark, achieving higher accuracy with far fewer training tokens. In a knowledge‑editing setting (AToKE), SD‑DPO attains an overall accuracy of 0.982, correctly answering queries with either new or old facts based on the requested time period.

By Takayuki Yamamoto, Daisuke Kawahara
arXiv AI
Sep 1

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process

The paper introduces LLM‑PeerReview, an unsupervised ensemble method that selects the best response from multiple LLM-generated candidates by scoring each answer with several LLMs, aggregating those scores via averaging or a graphical model, and choosing the highest-scoring response. The approach is peer‑review inspired, transparent, and interpretable, and it outperforms the Smoothie‑Global model by 6.9%–7.3% across factual recall QA, math reasoning, and instruction‑following tasks. The authors also provide a curated benchmark suite of 12 ensemble methods evaluated on four datasets and three task families to aid reproducibility.

By Zhijun Chen, Zeyu Ji, Qianren Mao, Hao Wu, Jinhuan Song, Junhang Cheng, Bangjie Qin, Zhuoran Li, Jingzheng Li, Kai Sun, Zizhe Wang, Yikun Ban, Zhu Sun, Xiangyang Ji, Hailong Sun, Xiao Huang
arXiv AI
Aug 25

LLM-Specific Utility for Retrieval-Augmented Generation

The paper introduces the concept of LLM‑specific utility, defining it as the performance gain a target large language model (LLM) achieves when provided with a passage compared to answering without evidence. A benchmark of utilitarian passages is built for four LLMs (Qwen3‑8B/14B/32B and Llama 3.1‑8B) across three QA datasets, revealing that each model benefits most from its own tailored evidence and that evidence optimized for other models is consistently suboptimal. The authors also create SpecUBench, a benchmark for LLM‑specific utility judgment, and show that current utility‑aware retrieval methods largely capture model‑agnostic usefulness, struggling to estimate LLM‑specific utility. "whyItMatters":"The study demonstrates that retrieval‑augmented generation must consider model‑specific evidence selection to truly improve LLM performance, highlighting a gap in existing utility‑aware methods."

By Hengran Zhang, Keping Bi, Jiafeng Guo, Jiaming Zhang, Shuaiqiang Wang, Dawei Yin, Xueqi Cheng