Natural language processing

Classical and neural NLP: translation, question answering, tokenization and the evaluation of language understanding.

2,641 stories · RSS feed

arXiv Computation and Language
Aug 25

LuxIT: A Luxembourgish Instruction Tuning Dataset from Monolingual Seed Data

LuxIT is a monolingual instruction‑tuning dataset for Luxembourgish, created by synthesizing instruction‑answer pairs from native texts using the DeepSeek‑R1‑0528 model and a quality‑assurance LLM‑as‑judge process. The resulting 227,507 high‑quality pairs were used to fine‑tune 14 LLMs (≤15 B parameters), yielding an average accuracy increase of +5.37 percentage points on standardized Luxembourgish proficiency exams and improvements in macro‑averaged F1 on nine of the fourteen downstream NLP tasks. These findings demonstrate that synthetic monolingual data can effectively enhance LLM performance in low‑resource languages and reveal the complex relationship between exam performance and practical NLP gains.

By Julian Valline, Cedric Lothritz, Siwen Guo, Jordi Cabot
arXiv Computation and Language
Aug 25

Mechanistic Interpretability of Chain-of-Thought Reasoning via Sequential Activation Patching

The paper introduces a sequential activation patching framework to study how Chain-of-Thought (CoT) prompting influences large language models over multiple generated tokens. By tracking CoT-conditioned attention-head activations across token positions and aggregating them with Part-of-Speech guidance, the authors identify distributed head sets that jointly contribute to answer generation. Targeted zero-ablation experiments confirm that these heads are functionally important, affecting mechanisms such as reasoning-trajectory maintenance, answer anchoring, exemplar-target separation, and numerical generation.

By Murat Dura, Serkan \"Ozt\"urk, Selma Tekir
arXiv Computer Vision
Aug 25

VQ-Transplant: Efficient VQ-Module Integration for Pre-trained Visual Tokenizers

VQ-Transplant is a framework that allows new vector‑quantization (VQ) modules to be inserted into frozen, pre‑trained visual tokenizers without retraining the entire model. By preserving all encoder‑decoder parameters and adding a lightweight decoder adaptation trained for only five epochs on ImageNet‑1k, the method mitigates decoder‑quantization mismatch. Experiments show that VQ-Transplant achieves near state‑of‑the‑art reconstruction fidelity for industry‑level models such as VAR while cutting training costs by 95%.

By Xianghong Fang, Yuan Yuan, Dehan Kong, Tim G. J. Rudner
arXiv Computer Vision
Aug 25

An end-to-end-trained vision-language model for native-language prostate pathology report generation

An end-to-end-trained vision-language model generates prostate biopsy reports in native languages, demonstrated in German. The system uses a tokenizer and model trained from scratch and an automated pipeline that splits composite reports into image-text pairs, producing 17,344 pairs from 2,402 cases without manual annotation. Evaluated on clinical attributes, it achieves 96.2% F1 for malignancy detection and 65.2% for Gleason grading, comparable to an FDA-cleared classifier and validated on external cohorts.

By Christian Grashei, Fabian G\"ulhan, Maximilian Legnar, Fabian St\"ogbauer, Cleo-Aron Weis, Carolin Mogler, Peter Sch\"uffler
arXiv AI
Aug 25

OVIBench: Benchmarking Online Video Question Answering under Interruption

OVIBench introduces the first standardized benchmark for evaluating vision‑language models on Online Video Question Answering under Interruption, a realistic setting where users can interrupt the model during answer generation. The benchmark categorizes interruptions into Cancellation, False Trigger, and Correction, supports both open‑ended and multiple‑choice tasks, and provides an offline simulation protocol plus a multi‑dimensional metric suite. Experiments show that OVIBench can distinguish models’ interruption‑handling abilities, particularly in following correction requests, and that fine‑tuning on the newly created OVI‑Train dataset yields significant performance gains.

By Naiming Liu, Zhiheng Wu, Shuning Wang, Tie Zhang, Bowen Liu, Tong Wang
arXiv AI
Aug 25

AraDetox: A Multi-Dialect Arabic Detoxification Dataset

AraDetox is a newly released multi-dialect Arabic detoxification dataset containing 10,500 harmful social‑media posts and 84,000 detoxified rewrites generated by GPT‑5 and Gemini 2.5 Flash across Modern Standard Arabic, Gulf, Levantine, and Egyptian Arabic. Human evaluation and automatic analyses confirm that the rewrites effectively remove harmful language while preserving meaning, lexical change, and dialectal style. The dataset is publicly available to support future research in Arabic detoxification, safe text generation, and multi‑dialect NLP.

By Mo El-Haj
arXiv AI
Aug 25

ReasonEdit: Editing Vision-Language Models using Human Reasoning

ReasonEdit is a new editor for vision‑language models that allows users to provide reasoning explanations during the editing process. It stores human reasoning in a codebook and retrieves relevant facts at inference time using a topology‑balanced multimodal embedding approach inspired by network science. Experiments on four VLMs and multiple rationale‑based visual question answering datasets show that incorporating human reasoning leads to state‑of‑the‑art editing performance and better generalization.

By Jiaxing Qiu, Kaihua Hou, Roxana Daneshjou, Ahmed Alaa, Thomas Hartvigsen
arXiv Computation and Language
Aug 25

Ask, Condition or Abstain: Reinforcement Learning for Missing-Premise Reasoning

The paper introduces Ask-Condition-Abstain Reinforcement Learning (ACA‑RL), a framework that trains reasoning models to handle queries missing a premise by either asking for it, conditioning on the unknown, or abstaining. ACA‑RL uses a reasoning‑graph‑guided pipeline to generate training instances with localized gap annotations and a structured reward over five observable response behaviors. The authors also present the Missing‑Premise Benchmark (MPB), a 274‑instance, human‑verified dataset covering mathematical, logical, and real‑world word problems, and show that ACA‑RL improves performance on MPB while maintaining competitive results on well‑posed tasks for Qwen3 and Llama models.

By Yongqi Tong, Zhenyu Zhang, Zimi Liu, Kewei Fu, Mingli Song, Haofei Zhang, Junshao Zhang, Hong Zhu, Jiang-Ming Yang, Xin Zhang, Jianshe Li