The Order Matters: Sequential Fine-Tuning of LLaMA for Coherent Automated Essay Scoring
arXiv:2606. 10327v1 Announce Type: cross Abstract: Automated Essay Scoring (AES) systems must judge interdependent discourse elements (e.
The Writerslogic team participated in the CLEF 2026 SimpleText shared task, tackling both text simplification (Task 1) and complexity spotting (Task 2). For simplification, they built a multi‑candidate pipeline with GPT‑4o‑mini, selecting the best candidate via a reference‑free heuristic, and their Claude Sonnet 4 submission achieved a SARI of 47.43 and BLEU of 14.21, ranking third overall on the Task 1 leaderboard. For complexity spotting, they fine‑tuned a DeBERTa‑v3‑large NLI model on 350 K labeled pairs, achieving a macro F1 of 0.8081 (0.8085 in an ensemble) on binary over‑generation identification and 0.804 accuracy on multi‑class error classification, placing them second among unique teams.
arXiv:2606. 10327v1 Announce Type: cross Abstract: Automated Essay Scoring (AES) systems must judge interdependent discourse elements (e.
arXiv:2606. 24259v1 Announce Type: cross Abstract: Fine-tuned encoders deployed across heterogeneous NLP tasks face three compounding problems: mismatched inductive biases, class-imbalance corruption of feature statistics, and no mechanism to condition attention on external lexical knowledge.
The paper evaluates two large language models, Claude Sonnet 4.5 and Claude Opus 5, on the bidirectional English Resource Grammar (ERG) tasks of generating English from Minimal Recursion Semantics (MRS) and parsing English into MRS. In generation, Opus achieves 76.3 BLEU—surpassing a 72k‑pair trained system and matching a million‑pair system—while Sonnet scores 65.7 BLEU, rising to 69.6 when selecting from ACE’s candidates. In parsing, both models lag behind ACE, attaining only 57.2 and 65.5 F₁ respectively, with exact‑match on about 1 % of sentences, highlighting that high generation scores do not guarantee accurate semantic parsing.
arXiv:2609.13154v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have made prompts increasingly large and complex. Techniques such as chain-of-thought reasoning (Wei et...
arXiv:2607. 04223v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) reduces but does not eliminate hallucination, and existing detectors return a single answer-level score that does not indicate which sentence is unsupported, or why.
arXiv:2607. 25675v1 Announce Type: new Abstract: Text-space optimization adapts large language models (LLMs) by editing external natural-language artifacts rather than model weights, so the optimized artifacts remain inspectable and the model can be treated as a black box.
The paper proposes a training‑free detector that uses sentence‑level context sensitivity to identify unsupported content in retrieval‑augmented generation (RAG) answers. By re‑scoring each sentence with full context, no context, and each chunk removed, the method flags the chunk whose removal most reduces a sentence’s likelihood as the likely source. Evaluated on RAGTruth, TofuEval, and RAGBench, the detector outperforms answer‑level faithfulness scores, achieving AUCs up to 0.745 and matching per‑chunk fact‑checkers while using only a fraction of the compute required by large‑language‑model judges.
The paper introduces Corpus Task Complexity (CTC), a metric that captures how a task’s difficulty scales with corpus size. It distinguishes low‑CTC tasks, whose difficulty grows linearly, from high‑CTC tasks, whose difficulty grows quadratically or more, and presents ten new high‑CTC tasks. Experiments show that models performing well on low‑CTC tasks often fail on high‑CTC tasks, highlighting the need for new approaches to large‑corpus reasoning.
arXiv:2608. 03803v1 Announce Type: cross Abstract: Multilingual language models are deployed across a hundred or more languages, yet most benchmarks test whether a model can perform a task _in_ a language rather than whether it commands the language itself, conflating fluency with proficiency.
The paper introduces a multi‑signal pipeline for detecting hallucinations in large language models, combining fine‑tuned DeBERTa‑v3 classification, Monte Carlo Dropout uncertainty, and temperature‑scaled calibration. On the HaluEval benchmark it achieves high performance (F1 = 0.915, AUROC = 0.977) across QA, summarization, and dialogue, and shows that 25 % of training data yields 77 % of full‑data performance. The authors also demonstrate that applying Direct Preference Optimization to a Qwen2.5‑0.5B generator cuts hallucination rates from 85.5 % to 37.7 %, and that domain‑specific fine‑tuning (PubMedBERT on SciFact) outperforms general‑domain models for biomedical text.
The paper presents the Writerslogic systems for three PAN 2026 shared tasks—Reasoning Trajectory Detection, Voight‑Kampff Generative AI Detection, and Multi‑Author Writing Style Analysis—using a unified analytical framework that prioritizes feature robustness under distribution shift. The framework distinguishes domain‑anchored, domain‑portable, and domain‑invariant features, explaining why generator‑specific traits fail while vocabulary fingerprints, compression measures, and character n‑grams remain effective. The authors report first‑place source detection and third‑place safety classification on Reasoning Trajectory Detection, a top‑scoring ensemble for Voight‑Kampff, and a detailed design for Multi‑Author Writing Style Analysis that was not evaluated due to a platform mix‑up.
arXiv:2606. 29809v1 Announce Type: cross Abstract: Hallucination detection has become a pressing requirement for trustworthy AI deployment at scale.