arXiv AI

Where LLM Graders Succeed and Break: Evidence from Two Computer-Science Exams

The study evaluates large language model (LLM) graders on two computer‑science exams, testing 171 configurations of closed‑ and open‑weights models. While the best LLM configuration achieved a mean absolute error of 1.64/35—better than the 2.61/35 error between two human graders—its performance was highly sensitive to the prompt. A short "strict grader" preamble caused most open‑weight models to exceed acceptable error thresholds or stop grading entirely, whereas fine‑tuning with a single LoRA adapter restored parity with human graders and reduced sensitivity to harsh prompts.

arXiv AI
Aug 19

Grading Needs a Rubric, Not Intelligence

Small language models can grade open‑ended exam answers as reliably as much larger models when they use an explicit rubric. In experiments with six cost‑efficient model configurations, the rubric decouples grading from judge intelligence, with answer identity explaining 95.6% of score variance and judge identity only 0.2%. Removing rubric criteria or the official answer collapses reliability and inflates scores, showing the rubric’s essential role.

By Jhen-Ke Lin
arXiv AI
Aug 20

Measuring the Partial-Credit Gap: A Strict Benchmark on Vietnam's 2025 Convex Marking Scheme

The paper introduces THPT‑Ladder, a benchmark based on Vietnam’s 2025 National High School Graduation Examination’s convex grading scheme, which rewards partial credit non‑additively. It shows that standard accuracy metrics inflate model scores because they treat partial knowledge proportionally, whereas the official rubric penalizes incomplete correct sets. Using the benchmark, the authors demonstrate that this discrepancy can shift a model’s percentile ranking by up to 13 points among over 480,000 candidates.

By Nguyen Quoc Hung, Nguyen Dang Minh, Le Nhu Quynh, Tran Khanh Linh, Nguyen Kieu Linh
arXiv Machine Learning
Sep 18

How Far Can Sub-3B Open Language Models Go in Zero-Shot Essay Scoring on an 8 GB Consumer GPU?

The study evaluates zero‑shot essay scoring using sub‑3B open language models that run locally on a single 8 GB consumer GPU. Four instruction‑tuned models (Qwen2.5‑0.5B, 1.5B, 3B and SmolLM2‑1.7B) were tested on all eight ASAP‑AES prompts, comparing rubric‑decomposed versus holistic prompting, different aggregation methods, and trait‑mapping strategies. Results show rubric‑decomposed prompting consistently outperforms holistic prompting, trait‑mapping is sensitive to calibration, and longer essays reduce error, yet the best local configuration (macro QWK 0.388) still falls short of human agreement and a length‑only baseline, suggesting these models are best suited for formative, human‑supervised feedback.

By Nguyen Dung Son, Dang Quang Minh, Nguyen Huu Loi, Truong Viet Vu, Nguyen Thai Anh
arXiv AI
Jul 17

Answer-Conditioned Chains of Thought Degrade Verifiable-Reasoning Distillation in Large Language Models

arXiv:2607. 14552v1 Announce Type: cross Abstract: A standard recipe for distilling the reasoning ability of large language models (LLMs) is to sample chains of thought from the model, keep those that reach the correct final answer, and fine-tune on the survivors.

By Jungseob Lee, Seungyoon Lee, Suhyune Son, Dongyub Jude Lee, Sungbin Han, Sugyeong Eo, Heuiseok Lim
arXiv Machine Learning
2d ago

Unknown-Traffic Detection, Calibration and Shortcut Reliance in Distilled Encrypted-Traffic Classifiers over One Year

The study investigates what knowledge a student model inherits from its teachers beyond accuracy when using knowledge distillation for encrypted‑traffic classification. By distilling a 101k‑parameter student from two teachers of equal accuracy but different construction, the authors test ten hypotheses over a year of real TLS traffic, finding that unknown‑traffic detection and shortcut reliance can transfer depending on temperature settings and model size, while other abilities do not. The results show that distillation can propagate teacher habits, but some inherited capabilities can also be achieved without a teacher.

By Mahmoud Abbasi
arXiv Machine Learning
Sep 15

GRADE: Graph Representation of LLM Agent Dependency and Execution

The paper introduces GRADE, a graph-based representation of large language model (LLM) agent executions that captures both execution steps and their dependencies. By adding graded dependency edges—observed, declared, or inferred—to the trace, the authors evaluate how this dependency layer affects failure prediction across six corpora involving tool use, coding, and web tasks. Experiments show that the dependency block can improve prediction in some settings, but its effectiveness varies with the evaluation probe and corpus, and controlled experiments demonstrate that the observed structure is not merely a degree-matched artifact.

By Yue Zhao