arXiv Computation and Language

A Layered Taxonomy for Chinese Learner Grammatical Error Annotation

The paper introduces a layered taxonomy for annotating grammatical errors in Chinese learner writing, aiming for consistency and linguistic relevance. It first classifies character- and punctuation-level orthographic errors by edit operation and subtype, then assigns other errors a three-layer core label that combines edit operation, linguistic domain, and part of speech, with optional Chinese-specific extensions. The taxonomy is evaluated through coverage analysis of automatically extracted edits and a preliminary consistency study using five large language models, confirming the layered approach while highlighting areas needing refinement.

arXiv AI
Jun 2

CSRP: Chain-of-Thought Reasoning for Chinese Text Correction via Reinforcement Learning with Efficiency-Aware Rewards

arXiv:2606. 00020v1 Announce Type: cross Abstract: Large Language Model (LLM) based Chinese Grammatical Error Correction (CGEC) systems face two critical challenges: general-purpose models lack specialized linguistic priors for subtle grammatical distinctions, and Supervised Fine-Tuning (SFT) with Maximum Likelihood Estimation fails to optimize for precision-focused metrics, leading to systematic over-correction.

By Wei Tian, Yuhao Zhou, Man Lan
arXiv Computation and Language
Aug 27

GUIDE: Generative Unsupervised Chinese Query Correction via Phonetic and Visual Shared-ID Encoding

The paper introduces GUIDE, a generative unsupervised framework for Chinese query correction that uses shared-ID encoding for phonetically or visually confusable characters and an encoder–decoder architecture to reconstruct queries within plausible confusion neighborhoods. It incorporates a time‑decayed, query‑frequency‑weighted objective to adapt to rapidly changing query vocabularies. Experiments on QSpell 250K and the large‑scale KwaiSearch dataset demonstrate that GUIDE consistently outperforms strong baselines, with online A/B testing confirming improvements in correction quality and downstream engagement.

By Lei Yang, Binbin Huang, Jiwei Tan, Xuhui Sui, Chang Tu, Yi Wang, Han Li
arXiv Computation and Language
Sep 4

Contextual Tamil Spelling and Grammar Correction Using Progressively Fine-Tuned Sequence-to-Sequence Transformers

The paper presents an end‑to‑end sequence‑to‑sequence approach for correcting Tamil spelling and grammar errors, leveraging progressively fine‑tuned transformer models (mT5‑small and mBART‑50). Using a synthetic corpus of 657,720 noisy‑clean sentence pairs across ten error categories, the authors introduce a four‑stage training schedule that targets surface noise, contextual grammar, single‑site sandhi, and multi‑site cross‑word sandhi. The best model, mBART‑50 v5, achieves 69.3% exact‑match accuracy on a balanced diagnostic set, with notable gains in sandhi (87.5%) and subject‑verb agreement (43.5%) accuracy, while also revealing a precision‑recall trade‑off for sandhi corrections.

By Karthikeyan A, Jaya Nirmala S, Sangeetha Sivanesan, Indhu R, Pranav Kumar, Bharat Jude Johnson, Vishnu Ram
arXiv AI
Jun 3

Fixing FOLIO and MALLS: Verified Annotations and an LLM-assisted Framework to Focus Human Relabeling

arXiv:2606. 02837v1 Announce Type: cross Abstract: Accurate translation from Natural Language to First-Order Logic (NL-to-FOL) underpins neurosymbolic AI systems and Natural Language Inference (NLI), making the quality of NL-to-FOL benchmarks essential -- yet these datasets have never been rigorously audited.

By Andrea Brunello, Cristian Curaba, Luca Geatti, Michele Mignani, Angelo Montanari, Nicola Saccomanno
arXiv AI
Jul 14

Automated Textbook Auditing with Multi-Agent LLM Systems

arXiv:2607. 11276v1 Announce Type: cross Abstract: Ensuring the quality of educational materials requires more than standard proofreading: textbooks must be audited for factual accuracy, domain-specific technical correctness, and linguistic quality simultaneously -- a task that general-purpose grammar checkers cannot address.

By Ciprian Cristescu, Adrian-Marius Dumitran, Angela-Liliana Dumitran, Gabriel Stefan