The Order Matters: Sequential Fine-Tuning of LLaMA for Coherent Automated Essay Scoring
arXiv:2606. 10327v1 Announce Type: cross Abstract: Automated Essay Scoring (AES) systems must judge interdependent discourse elements (e.
Edu-QuRating is a pipeline that scores and curates educational data across multiple dimensions—accuracy, engagement, structure, and audience appropriateness—using an LLM judge to label document pairs and distill these preferences into reusable Edu-QuRaters. The best Edu-QuRater achieves 91.7% accuracy against held‑out GPT‑4.1‑mini judgments and is applied to filter 322.25 M FineWeb‑Edu‑Fortified documents, improving small‑model pre‑training performance on nine benchmarks. Additionally, Edu-QuRater scores serve as reward signals in GRPO post‑training, yielding responses that are preferred for pedagogical quality and instruction following over the Qwen3‑4B base model.
arXiv:2606. 10327v1 Announce Type: cross Abstract: Automated Essay Scoring (AES) systems must judge interdependent discourse elements (e.
arXiv:2606. 06546v1 Announce Type: new Abstract: Evaluating large language models (LLMs) for education requires measuring how models teach, not only what they know.
The study evaluates how well pre‑trained models can classify the Bloom level of AI‑generated educational questions, a task that is crucial for ensuring pedagogical quality. Traditional machine‑learning models perform poorly on out‑of‑distribution data, whereas transformer and large‑language models achieve higher accuracy, especially after feature‑engineering techniques such as text splicing and appending learning objectives. Retraining the models yields the most significant performance gains across all datasets.
AI University (AI‑U) is a flexible framework that uses a fine‑tuned large language model (LLM) combined with retrieval‑augmented generation (RAG) and a reasoning synthesis model to produce style‑aligned responses from lecture videos, notes, and textbooks. In a graduate‑level finite‑element‑method (FEM) course, the authors created a pipeline to generate course‑grounded training data, fine‑tuned an open‑source LLM with Low‑Rank Adaptation (LoRA), and applied RAG‑based synthesis. Evaluation through cosine similarity, LLM‑based assessment, expert review, and user studies showed that the expert model outperformed the base model in alignment with course materials, with 86 % of test cases scoring higher and human users preferring the expert model roughly twice as often. whyItMatters":"The study demonstrates a practical method for building course‑specific learning assistants that improve alignment with instructional content, offering a template that can be extended across STEM fields."
E-CONAN introduces Arabic textual entailment and natural inference benchmarks comprising two datasets: E-CONAN-2 (2-way RTE) and E-CONAN-3 (3-way NLI). The datasets are built from automatically-translated pairs, human-validated machine translations, hand-crafted pairs from Arabic teaching books, and rumor-containing news headlines. The authors evaluated nine multilingual pretrained models and five large language models on these benchmarks, demonstrating that E-CONAN offers a more diverse and robust assessment than existing datasets like XNLI and ArNLI.
arXiv:2509. 16780v3 Announce Type: replace-cross Abstract: Large language models (LLMs) show promise as educational aids but often lack alignment with specific course materials.
The paper introduces Style‑Debiased DPO (SD‑DPO), a method that refines large language models’ ability to retrieve stored knowledge by using preference optimization that corrects for style differences while preserving factual accuracy. SD‑DPO evaluates on the EntiGraph storing‑side framework and outperforms baseline CPT on the QuALITY reading‑comprehension benchmark, achieving higher accuracy with far fewer training tokens. In a knowledge‑editing setting (AToKE), SD‑DPO attains an overall accuracy of 0.982, correctly answering queries with either new or old facts based on the requested time period.
arXiv:2510. 06048v4 Announce Type: replace Abstract: Effective data selection is essential for pretraining large language models (LLMs), enhancing efficiency and improving generalization to downstream tasks.
arXiv:2609.23088v1 Announce Type: new Abstract: Educational foundation models must solve problems, understand curriculum structure, diagnose learner difficulties, and provide appropriate instructiona...
arXiv:2609.08797v2 Announce Type: replace-cross Abstract: The Potts model extends the Ising model to multinomial data. We introduce a Rater Ising-Potts model that uses agreement indicators between pa...
The paper introduces LLM‑PeerReview, an unsupervised ensemble method that selects the best response from multiple LLM-generated candidates by scoring each answer with several LLMs, aggregating those scores via averaging or a graphical model, and choosing the highest-scoring response. The approach is peer‑review inspired, transparent, and interpretable, and it outperforms the Smoothie‑Global model by 6.9%–7.3% across factual recall QA, math reasoning, and instruction‑following tasks. The authors also provide a curated benchmark suite of 12 ensemble methods evaluated on four datasets and three task families to aid reproducibility.
The paper introduces the concept of LLM‑specific utility, defining it as the performance gain a target large language model (LLM) achieves when provided with a passage compared to answering without evidence. A benchmark of utilitarian passages is built for four LLMs (Qwen3‑8B/14B/32B and Llama 3.1‑8B) across three QA datasets, revealing that each model benefits most from its own tailored evidence and that evidence optimized for other models is consistently suboptimal. The authors also create SpecUBench, a benchmark for LLM‑specific utility judgment, and show that current utility‑aware retrieval methods largely capture model‑agnostic usefulness, struggling to estimate LLM‑specific utility. "whyItMatters":"The study demonstrates that retrieval‑augmented generation must consider model‑specific evidence selection to truly improve LLM performance, highlighting a gap in existing utility‑aware methods."