arXiv:2606. 01982v1 Announce Type: new Abstract: Schema-constrained information extraction from diverse educational and labor-market corpora remains an open challenge in natural language processing because existing pipelines rely primarily on lexical-surface methods that cannot recover implicit competencies, lack grounding in shared taxonomies, and provide no formal measures of extraction reliability or document-level completeness.
By Sherzod Turaev, Mary John, Mamoun Awad, Nazar Zaki, Khaled Shuaib
arXiv:2608. 12356v1 Announce Type: cross Abstract: A college offering several overlapping computing degrees implicitly assumes that its programs are differentiated in line with how the labor market segments computing work and that, together, they prepare graduates for that market.
By Sherzod Turaev, Saja Aldabet, Mary John, Namya Musthafa, Mamoun Awad, Nazar Zaki, Khaled Shuaib
arXiv:2606. 15349v1 Announce Type: cross Abstract: Standardized examinations are typically treated as uniform syllabus coverage problems.
By Joy Bose, Om Thomas
CaSKG introduces a counterfactual‑causal skill graph framework that calibrates procedural relations before retrieval, building a high‑recall directed candidate graph from semantic, lexical, input/output, and structural evidence and refining it with repair evidence and optional LLM judgment. The framework applies direction‑conditioned textual counterfactual probes—removing, substituting, and reordering skill pairs—to aggregate evidence with Bayesian smoothing, producing a state‑filtered weighted graph for task‑conditioned expansion. Evaluated across six LLM backbones on ALFWorld and ScienceWorld, CaSKG outperforms existing Graph‑of‑Skills methods, improving macro‑average scores and reducing mean environment steps while preserving essential skill dependencies.
By Zhiyuan Li, Linyuan Gao, Xuechun Ding, Hongwei Chen, Yuan Wu, Yi Chang
arXiv:2607. 02432v1 Announce Type: new Abstract: Scalable and reliable grading of command-line examinations remains a challenge in computing education, where rising enrolments make manual marking difficult and rule-based autograders cannot handle partial credit, equivalent solutions, or syntactic variation.
By Manuel Alonso-Carracedo, Ruben Fernandez-Boullon, Pedro Celard, Francisco J. Rodriguez-Martinez, Lorena Otero-Cerdeira
The paper introduces Intelligent Target Locator (ITL), a method that measures how well a document aligns with concepts in a Structured Reference Document (SRD) by creating concept‑specific term profiles and computing a textual‑unit–concept affinity matrix. ITL assigns importance weights to terms based on concept membership, term specificity, and discriminability, enabling traceable, quantitative alignment scores at multiple granularity levels. An internal consistency test on the 17 Sustainable Development Goals showed that each goal statement achieved its highest affinity with its corresponding concept, demonstrating ITL’s ability to distinguish conceptual profiles.
By Ra\'ul Gir\'aldez, Dayrelis Mena, Jes\'us S. Aguilar--Ruiz
arXiv:2607. 00140v1 Announce Type: cross Abstract: As computing education expands beyond traditional programming into operational domains such as systems administration and command-line environments, existing pedagogical frameworks struggle to capture a dimension that is critical in these contexts: the real-world consequences of learner actions.
By Manuel Alonso-Carracedo (Universidade de Vigo, Spain, IFCAE, Universidade de Vigo, Spain), Ruben Fernandez-Boullon (Universidade de Vigo, Spain, IFCAE, Universidade de Vigo, Spain), Pedro Celard (Universidade de Vigo, Spain, IFCAE, Universidade de Vigo, Spain), Francisco J. Rodriguez-Martinez (Universidade de Vigo, Spain, IFCAE, Universidade de Vigo, Spain), Lorena Otero-Cerdeira (Universidade de Vigo, Spain, IFCAE, Universidade de Vigo, Spain)
arXiv:2608. 13708v1 Announce Type: cross Abstract: Automatically generating textbook-grounded assessment items can reduce science teachers' workload, but existing retrieval-augmented generation (RAG) systems rely on flat retrieval, support only single-question generation, lack safeguards against weak evidence, and are ill-suited to low-resource, board-exam-structured curricula.
By Fatema Tuj Johora Faria, Mukaffi Bin Moin, M. F. Mridha, Jubayer Al Mahmud
The study evaluates six retrieval methods for ranking computer science faculty as potential academic advisors based on graduate applicants’ research interest statements. Using a new dataset of 768 faculty profiles from nine U.S. universities and 162 graded relevance judgments across five queries, the reranked hybrid approach achieved the highest mean NDCG@10 (0.477). Ablation experiments showed that faculty biographies alone outperform the full model, and adding arXiv abstracts actually decreased performance, leading to a late‑fusion design. All code, data, and labels are publicly released.
By Biraj Subedi
arXiv:2607. 14707v1 Announce Type: cross Abstract: Large language models routinely produce fluent answers to single-shot prompts, yet deploying them as reliable components of a domain decision system is substantially harder.
By Akash Raj
arXiv:2606. 06546v1 Announce Type: new Abstract: Evaluating large language models (LLMs) for education requires measuring how models teach, not only what they know.
By Tao Liu, Ye Lu, Ruohua Zhang, Siyu Song, Wentao Liu, Aimin Zhou, Hao Hao
The paper introduces a structured approach to extracting skill and responsibility level pairs from free text using the Skills Framework for the Information Age (SFIA). It evaluates five methods—including lexical baselines, retrieval‑augmented generation, and multi‑agent crews—against expert‑mapped European ICT role profiles, finding that generative strategies are more precise and that only explicit level‑prediction strategies reliably assign responsibility levels. The study also releases an automated SFIA‑9 corpus and establishes the first reproducible baseline for level‑aware skill extraction.
By Ranuga Disansa, U. S. Samarasinghe, Lasith Gunawardena