Natural language processing

Classical and neural NLP: translation, question answering, tokenization and the evaluation of language understanding.

2,564 stories · RSS feed

arXiv Computation and Language
Sep 23

ICDAR2026 Competition on Multimodal Reasoning over Documents in Multiple Domains

The ICDAR2026 Competition on Multimodal Reasoning over Documents in Multiple Domains introduced a new Visual Question Answering benchmark that tests reasoning over documents from eight distinct domains such as business reports, scientific papers, and engineering drawings. Twenty valid submissions from eight teams were evaluated, featuring approaches ranging from zero‑shot vision‑language models to multi‑agent ensembles and fine‑tuned multimodal systems. Results indicate that the most effective systems employ structured evidence extraction, retrieval, verification, and orchestration across multiple components rather than single‑pass prompting.

By Artemis Llabr\'es, Marc Serra Ortega, Tom\`as Ockier, Samuel Ortega Cuadra, Amritpal Singh, Christos Georgakilas, Andrey Barsky, Ernest Valveny, Dimosthenis Karatzas
arXiv Computation and Language
Sep 23

ufakzeka-1: Building and Evaluating a 151M-Parameter Turkish Language Model from Scratch

The paper introduces ufakzeka-1, a 151‑million‑parameter Turkish language model trained from scratch on 13.5 B tokens. It details the tokenizer, a three‑stage pretraining schedule, post‑training data augmentation, and a comprehensive evaluation suite that includes release gates, a rule‑checked conversation sweep, and hand tests—all with prompts excluded from training data. The authors report three key findings about safety gate performance, training‑seed variance, and data‑round effects, and they release the model weights, data recipe, evaluation code, and spend ledger under Apache‑2.0.

By Sait Furkan Teke (ufak AI)
arXiv Computation and Language
Sep 23

DTOC: Dynamic Tool Output Compression for Adaptive Context Management in AI Agents

The paper introduces Dynamic Tool Output Compression (DTOC), a framework that manages context in large‑language‑model agents by storing full tool outputs in external memory and inserting compact placeholders into the active context. DTOC treats context updates as explicit, reversible operations within the agent’s reasoning loop, allowing selective reconstruction of compressed outputs when needed. Experiments on the DeepSWE benchmark show that for responsive models such as Sonnet 4.6 and GPT‑5.4, DTOC reduces input tokens and agent steps while significantly improving solve rates and lowering cost per solved task, with ablation studies confirming the importance of reversibility for maintaining performance.

By Abhay Chaturvedi, Shreya Bhattacharya, Rashmika Gopalkrishnan, Peter van der Putten
arXiv Computation and Language
Sep 23

Explanation-Guided Medical Named Entity Recognition with Stability and Boundary Awareness for Atopic Dermatitis

The paper introduces a stability and boundary-aware explanation-guided framework for medical named entity recognition (NER) in Chinese atopic dermatitis clinical texts. It uses perturbation-based analysis to assess explanation stability and entity boundary sensitivity, and an adaptive fusion strategy to combine local and global explanations. The fused explanations are integrated into training via stability, boundary-aware, and consistency constraints, leading to improved recognition performance and more reliable explanations across multiple NER models.

By Xueguang Li (School of Information and Software Engineering, University of Electronic Science and Technology of China, Sichuan, China), Di Lin (School of Information and Software Engineering, University of Electronic Science and Technology of China, Sichuan, China), Xue Jiang (Department of Dermatology, Chongqing Traditional Chinese Medicine Hospital, Chongqing, China), Yanxi Li (Department of Dermatology, Chongqing Traditional Chinese Medicine Hospital, Chongqing, China), Yugang Chi (Chongqing Health Center for Women and Children, Chongqing, China)
arXiv Computation and Language
Sep 23

Domain-Adaptive Pretraining Enhances Water Treatment Semantic Representation for Large-Scale Structured Literature Mining

The paper introduces WaterBERT, a domain‑adapted encoder model trained on a 2.97‑billion‑token water treatment corpus to capture domain‑specific semantics for literature mining. Fine‑tuned versions of WaterBERT outperform general‑purpose and other domain BERT models on tasks such as treatment process classification, named entity recognition, and relation extraction. The authors also demonstrate WaterBERT’s utility in large‑scale processing, generating coherent research topics, building a structured knowledge graph from 693,211 abstracts, and creating a Water Knowledge‑Enhanced Retrieval System that surpasses text‑based baselines.

By Mudi Zhai (UNSW Water Research Centre, School of Civil and Environmental Engineering, The University of New South Wales, Sydney, NSW 2052, Australia), Ruihong Qiu (School of Electrical Engineering and Computer Science, The University of Queensland, Brisbane, QLD 4072, Australia), Qingyun Zeng (Microsoft Copilot Studio AI, Redmond, WA 98052, United States, Departments of Mathematics & Department of Computer and Information Science, University of Pennsylvania, Philadelphia, PA 19104, United States), T. David Waite (UNSW Water Research Centre, School of Civil and Environmental Engineering, The University of New South Wales, Sydney, NSW 2052, Australia), Bing-Jie Ni (UNSW Water Research Centre, School of Civil and Environmental Engineering, The University of New South Wales, Sydney, NSW 2052, Australia), Haoran Duan (UNSW Water Research Centre, School of Civil and Environmental Engineering, The University of New South Wales, Sydney, NSW 2052, Australia, Department of Civil Engineering, The University of Hong Kong, Pokfulam, Hong Kong SAR, China)
arXiv Computation and Language
Sep 23

Same Chart, Different Story: Bias in Vision-Language Chart Interpretation

The paper introduces ChartBias, a benchmark of 820 real-world charts covering six social attributes, designed to audit bias in vision‑language models (VLMs) that interpret charts. Across 12 VLMs, the study identifies three failure modes—narrative shift, group hallucination, and preference polarity—where models produce different or misleading narratives when the referenced social group changes. A multi‑agent mitigation framework is proposed, separating evidence extraction from group‑conditioned generation and using a counterfactual judge, which reduces narrative shift while maintaining chart‑grounded reasoning.

By Mizanur Rahman, Huan Wu, Arash Asgari, Enamul Hoque Prince, Laleh Seyyed-Kalantari
arXiv Computer Vision
Sep 23

StableVQ: Practical Guidelines for Stable Vector-Quantized Tokenizer Training

StableVQ introduces practical guidelines to improve training stability for vector‑quantized tokenizers used in image generation models. It addresses instability caused by the entanglement of encoder–decoder and codebook training by proposing three techniques: Dynamic STE for the encoder, Region VQ Loss for the codebook, and a Decoupled Schedule for independent learning rates. Experiments on ImageNet show consistent gains in stability, codebook utilization, and reconstruction quality across various settings.

By Bao Tang, Jiahao Guo, Haoxiang Cao, Wenyu Liu, Changqian Yu, Kun Gai, Xinggang Wang