Natural language processing

Classical and neural NLP: translation, question answering, tokenization and the evaluation of language understanding.

2,601 stories · RSS feed

arXiv AI
Sep 2

H2Table: Hierarchical Hypergraph-Enhanced Large Language Models for Complex Table Reasoning

H2Table introduces a hierarchical hypergraph representation for complex tables, enabling a hypergraph encoder to capture semantic relationships between headers and cells. The framework uses learnable query vectors to extract structural embeddings for large language models. Experiments on the HiTab dataset show a 22.88% improvement over state‑of‑the‑art baselines on tables with four levels of nesting.

By Jia Ling, Yangfan Wang, Chen Tang, Haoming Tan, Yang Yang, Yi Guan, Jingchi Jiang
arXiv AI
Sep 2

RW-LoRA: Communication-Efficient Decentralized LoRA Fine-Tuning via Random Walks

RW-LoRA introduces a random‑walk approach to fine‑tune LoRA models in a decentralized setting, using a single model token that moves through the network and updates locally. This eliminates the need for global synchronization and reduces communication and computation costs compared to centralized or gossip‑based methods. The authors provide convergence guarantees for non‑convex objectives and demonstrate competitive performance on NLP tasks across various graph topologies.

By Xingran Chen, Rohit Bhagat, Ghadir Ayache, Rawad Bitar, Yanmin Gong, Salim El Rouayheb
arXiv AI
Sep 2

Is Human Annotation Necessary? Iterative MBR Distillation for Error Span Detection in Machine Translation

The paper introduces Iterative MBR Distillation for Error Span Detection (ESD) in machine translation, a self‑evolution framework that replaces human annotations with pseudo‑labels generated by a large language model. By iteratively applying Minimum Bayes Risk decoding, the method produces high‑quality error spans without costly human effort. Experiments on WMT Metrics Shared Task datasets show that models trained solely on these pseudo‑labels outperform both unadapted baselines and supervised models trained on human data at system and span levels, while keeping sentence‑level performance competitive.

By Boxuan Lyu, Haiyue Song, Zhi Qu
arXiv Computation and Language
Sep 2

How Correct Is Your Answer? A Semantic Correctness Framework for Open QA Evaluation

The paper introduces a semantic correctness taxonomy that categorizes open‑ended QA answers into eight ordered classes, distinguishing between correct, verbose, and hallucinated responses. It releases two datasets—CAP‑Correctness and CAP‑Statements—to support benchmark evaluation and NLI‑based training. The authors also propose CAP (Context‑Aware Precision), a reference‑based metric that scores question‑conditioned statements via bidirectional NLI and demonstrates superior performance under a monotonicity protocol.

By Elitsa Yotkova, Violeta Kastreva, Petar Velkov, Hristo Boyanov, Dimitar Dimitrov, Ivan Koychev, Preslav Nakov
arXiv AI
Sep 2

EM^2Mem: Event-Centric Multimodal Memory for Large Language Models

arXiv:2609.00551v1 Announce Type: cross Abstract: Multimodal memory offers a scalable interface for long-video question answering, but existing methods often retrieve captions, frames, transcripts, s...

By Yijun Chen, Yaqi Zheng, Yanya Li, Boyi Xiao, Buqiang Xu, Shuofei Qiao, Jizhan Fang, Xinle Deng, Yunzhi Yao, Xuehai Wang, Liuxin Zhang, Hui Li, Huajun Chen, Shumin Deng
arXiv Computation and Language
Sep 2

When Tokenization is Secretly Output Supervision

The paper argues that tokenization in language models should be viewed as an output supervision decision rather than merely input preprocessing. In autoregressive models, the granularity of the tokenizer determines the supervision signal the model receives, influencing learning difficulty, internal representations, and task performance. Experiments on numeric reasoning show that output tokenization, rather than input tokenization, drives differences in performance and training dynamics, and a survey of recent CL papers reveals that tokenization choices are rarely reported or acknowledged.

By Tanja Baeumel, Josef van Genabith, Simon Ostermann
arXiv AI
Sep 2

Beyond Static Summarization: Proactive Memory Extraction for LLM Agents

The paper "Beyond Static Summarization: Proactive Memory Extraction for LLM Agents" identifies two shortcomings in current memory extraction for large language model agents: (1) extraction occurs ahead of time and mixes multiple types of information, leading to loss of useful details, and (2) extraction is typically one‑off, allowing errors and hallucinations to persist. To address these issues, the authors propose ProMem, a proactive framework that separates details, events, and relations, applies distinct extraction strategies for each, checks for completeness, and verifies facts at an atomic level. Experiments demonstrate that ProMem enhances memory completeness and question‑answering accuracy while maintaining a favorable balance between quality and token cost.

By Chengyuan Yang, Zequn Sun, Wei Wei, Wei Hu