arXiv AI

Automated grading of Linux/bash examinations using large language models: a four-level cognitive taxonomy approach

arXiv:2607. 02432v1 Announce Type: new Abstract: Scalable and reliable grading of command-line examinations remains a challenge in computing education, where rising enrolments make manual marking difficult and rule-based autograders cannot handle partial credit, equivalent solutions, or syntactic variation.

arXiv Machine Learning
Jul 2

CogTax: A Four-Level Cognitive Taxonomy for Command-Line Computing Education

arXiv:2607. 00140v1 Announce Type: cross Abstract: As computing education expands beyond traditional programming into operational domains such as systems administration and command-line environments, existing pedagogical frameworks struggle to capture a dimension that is critical in these contexts: the real-world consequences of learner actions.

By Manuel Alonso-Carracedo (Universidade de Vigo, Spain, IFCAE, Universidade de Vigo, Spain), Ruben Fernandez-Boullon (Universidade de Vigo, Spain, IFCAE, Universidade de Vigo, Spain), Pedro Celard (Universidade de Vigo, Spain, IFCAE, Universidade de Vigo, Spain), Francisco J. Rodriguez-Martinez (Universidade de Vigo, Spain, IFCAE, Universidade de Vigo, Spain), Lorena Otero-Cerdeira (Universidade de Vigo, Spain, IFCAE, Universidade de Vigo, Spain)
arXiv AI
Sep 7

A Human-in-the-Loop Framework for AI-Assisted Scoring in Large-Scale Writing Assessment

The paper introduces a human‑in‑the‑loop framework for AI‑assisted scoring of short written responses in a large‑scale national assessment. Using data from two recent test editions with about 5,000 student responses each, the authors validate that AI‑generated scores align moderately to highly with human raters across multiple rubric dimensions. The framework includes a correction workflow that flags cases needing human review, thereby reducing manual workload while maintaining assessment quality.

By Mar\'ia Eugenia Curi, Germ\'an Capdehourat, Isabel Amigo, Magdalena Romano, Rosana Serra, Adri\'an Silveira, Andr\'es Peri
Hugging Face Trending Papers
Sep 4

A Human-in-the-Loop Framework for AI-Assisted Scoring in Large-Scale Writing Assessment

The paper introduces a human‑in‑the‑loop framework for AI‑assisted scoring of short written responses in a large‑scale national assessment. Using data from two recent test editions with about 5,000 responses each, the study validates that AI-generated scores align moderately to highly with human raters across multiple rubric dimensions. The framework also identifies when human review is most needed, allowing more efficient allocation of expert effort while maintaining assessment quality.

arXiv AI
Aug 25

GIM: Evaluating models via tasks that integrate multiple cognitive domains

The paper introduces the Grounded Integration Measure (GIM), a benchmark of 820 expert‑authored problems designed to test models on tasks that integrate multiple cognitive operations such as constraint satisfaction, state tracking, epistemic vigilance, and audience calibration. GIM emphasizes realistic, broadly accessible knowledge rather than specialized expertise, and uses a judge‑aware 2‑parameter logistic IRT model to produce robust ability estimates across 53 model‑thinking‑level configurations. The authors provide a comprehensive leaderboard of 22 models and 47 test configurations, and conduct an extensive study on how test‑time compute affects model capability, finding that configuration choices like thinking budget and quantization can be as influential as model selection itself. whyItMatters":"By focusing on integration of multiple cognitive domains, GIM offers a more realistic assessment of model reasoning capabilities than benchmarks that either overemphasize memorization or abstract reasoning alone."

By Rohit Patel, Alexandre Rezende, Steven McClain
arXiv AI
Jun 19

Measuring Curriculum Alignment across Topical Coverage, Competency, and Cognitive Depth: A Longitudinal Framework Applied to CS2013 and CS2023

arXiv:2606. 19469v1 Announce Type: new Abstract: Undergraduate computer science is governed by international curricular guidelines revised about once a decade, yet programs lack a reliable, reproducible way to measure how completely they cover the current guidelines and how that coverage shifts when the guidelines are restructured.

By Sherzod Turaev, Mary John, Saja Aldabet, Mamoun Awad, Nazar Zaki, Khaled Shuaib