arXiv Machine Learning

Does Continued Pretraining on a Learner Corpus Improve Automated Essay Scoring on English Proficiency Tests? Evidence from EFCAMDAT

The study examines whether domain‑adaptive continued pretraining (DAPT) on a learner‑writing corpus (EFCAMDAT) can enhance transformer‑based automated essay scoring (AES) for English proficiency tests. Researchers applied DAPT to BERT, RoBERTa, and DistilBERT and compared the adapted models with their original checkpoints on the FCE and IELTS datasets, evaluating both in‑domain scoring and few‑shot cross‑dataset transfer. Results show that full‑corpus DAPT yields mixed effects, while proficiency‑specific DAPT often outperforms full‑corpus DAPT and sometimes even the non‑adapted baseline, though benefits vary by proficiency composition and encoder architecture and do not consistently transfer across tests.

arXiv Machine Learning
Aug 5

M-GATE: Multilingual Grammar, Accuracy in Translation, and Efficiency Benchmark for Large Language Models

arXiv:2608. 03803v1 Announce Type: cross Abstract: Multilingual language models are deployed across a hundred or more languages, yet most benchmarks test whether a model can perform a task _in_ a language rather than whether it commands the language itself, conflating fluency with proficiency.

By Tom\'a\v{s} Burkert, Angelika Peljak-{\L}api\'nska, David Zelen\'y
arXiv Computation and Language
Aug 25

LuxIT: A Luxembourgish Instruction Tuning Dataset from Monolingual Seed Data

LuxIT is a monolingual instruction‑tuning dataset for Luxembourgish, created by synthesizing instruction‑answer pairs from native texts using the DeepSeek‑R1‑0528 model and a quality‑assurance LLM‑as‑judge process. The resulting 227,507 high‑quality pairs were used to fine‑tune 14 LLMs (≤15 B parameters), yielding an average accuracy increase of +5.37 percentage points on standardized Luxembourgish proficiency exams and improvements in macro‑averaged F1 on nine of the fourteen downstream NLP tasks. These findings demonstrate that synthetic monolingual data can effectively enhance LLM performance in low‑resource languages and reveal the complex relationship between exam performance and practical NLP gains.

By Julian Valline, Cedric Lothritz, Siwen Guo, Jordi Cabot
arXiv Computation and Language
Aug 27

One Form to Transfer Them All: Pretraining Multilingual Language Models Beyond Native Orthography

The study evaluates how different input representations—orthographic text, IPA transcription, and romanization—affect cross‑lingual transfer in autoregressive multilingual language models. Across three model sizes and eight languages grouped into typologically motivated pairs, romanized pretraining consistently outperforms native orthography and IPA, especially as model scale increases. Fine‑tuning a text‑pretrained model on romanized data can harm performance on languages already covered by the base model, suggesting romanization should be integrated at pretraining rather than applied later.

By Muge Zhang, Aaron Jencks, Krishna Badikela, Yulia Tsvetkov, Sachin Kumar
arXiv Computation and Language
5d ago

Domain-specific Pretraining Profile and Transformer Performance: Evidence from Modeling Digital Pragmatics in Arabic-English Code-switching

arXiv:2609.14571v1 Announce Type: new Abstract: This study highlights the role of domain-specific pretraining profile (DSPP) in Transformer performance for modeling digital pragmatics in Arabic-Engli...

By Fahad Al Hussen, King Saud University, Riyadh, Saudi Arabia, Mohammed Q. Shormani, Ibb University, Ibb, Yemen
arXiv Computation and Language
Aug 27

Cross-Dataset Stability of Expert-Informed Skill Prompting and Fine-Tuning for Chinese Metaphor Identification

The study compares four approaches for Chinese sentence-level metaphor identification: BERT fine‑tuning, QLoRA-based large language model fine‑tuning, zero‑shot LLM prompting, and zero‑shot prompting with an expert‑informed procedural Skill. Results show that fine‑tuning yields the highest accuracy on the native test set, while the Skill‑based zero‑shot method provides the most stable performance across three datasets, achieving the highest external floor and the smallest performance range. Adding the Skill reduces false positives on one dataset but increases false negatives on others, indicating a trade‑off between precision and recall.

By Yufeng Wu, Meichun Liu
arXiv Machine Learning
6d ago

Limits of LLM Text Detectors in Education

The paper "Limits of LLM Text Detectors in Education" argues that existing LLM‑generated text detectors assume a binary human/LLM distinction, which fails to capture realistic student‑AI collaboration. It introduces a contribution‑aware evaluation framework with eight student contribution levels and presents GEDE, a benchmark of over 900 human‑written and 12,500 generated essays across 886 tasks. Using GEDE, the authors evaluate four detection methods and find that most detectors perform poorly on intermediate contribution levels, especially LLM‑assisted revisions, raising concerns about false accusations.

By Lukas Gehring, Benjamin Paa{\ss}en
arXiv AI
Aug 7

Bias Analysis of L2 Speaking Assessment Systems Using Concept Activation Vectors

arXiv:2608. 06300v1 Announce Type: new Abstract: Automatic speaking assessment systems are increasingly deployed in high-stakes settings to mark second language (L2) learners' speaking tests, making it critical to show that their scores depend on speaking proficiency rather than irrelevant speaker attributes such as first language (L1) or age.

By Arya Labroo, Mengjie Qian, Kate Knill
arXiv AI
Sep 7

A Human-in-the-Loop Framework for AI-Assisted Scoring in Large-Scale Writing Assessment

The paper introduces a human‑in‑the‑loop framework for AI‑assisted scoring of short written responses in a large‑scale national assessment. Using data from two recent test editions with about 5,000 student responses each, the authors validate that AI‑generated scores align moderately to highly with human raters across multiple rubric dimensions. The framework includes a correction workflow that flags cases needing human review, thereby reducing manual workload while maintaining assessment quality.

By Mar\'ia Eugenia Curi, Germ\'an Capdehourat, Isabel Amigo, Magdalena Romano, Rosana Serra, Adri\'an Silveira, Andr\'es Peri