Natural language processing

Classical and neural NLP: translation, question answering, tokenization and the evaluation of language understanding.

2,601 stories · RSS feed

arXiv Computation and Language
Sep 1

ManGo: Manga Active Narrative Grounding Optimization

ManGo is an unsupervised framework for manga visual question answering that actively selects panels, extracts concise clues, and decides when to stop, creating a compact evidence sketch before answering. It introduces Active Narrative Sketching (ANS) and optimizes its behavior using group-relative policy training with two rewards: answer preference from listwise self-ranking and path consistency from stable ordered panel trajectories. Experiments on standard manga understanding benchmarks demonstrate that ManGo achieves state‑of‑the‑art performance across different settings.

By Hao Qiu, Junyan Wang, Zheyuan Liu, Lei Fan, Hong Jia, Lianbo Guo, Zhulin Tao
arXiv Computation and Language
Aug 31

Long Story Short: Story-level Video Understanding from 20K Short Films

The paper introduces Short‑Films 20K (SF20K), a large publicly available movie dataset comprising 20,143 amateur films totaling 3,582 hours, with an average length of 12 minutes per film. Accompanying the dataset is SF20K‑Test, a manual open‑ended question‑answering benchmark featuring 95 movies and 979 question‑answer pairs. Analysis of the benchmark shows limited data leakage, highlights the necessity of long‑term reasoning, and demonstrates that instruction tuning on the large‑scale dataset significantly boosts vision‑language model performance.

By Ridouane Ghermi, Xi Wang, Vicky Kalogeiton, Ivan Laptev
arXiv Computation and Language
Aug 31

Select, Label, Evaluate: Active Testing in NLP

The paper introduces Active Testing, a framework that selects the most informative test samples for annotation in NLP, aiming to reduce human effort while accurately estimating model performance. Experiments across 18 datasets and 4 embedding strategies show up to 95% annotation savings with less than 1% loss in performance estimation accuracy. The authors also propose an adaptive stopping criterion to determine the optimal number of samples without a predefined budget.

By Antonio Purificato, Maria Sofia Bucarelli, Andrea Bacciu, Fabrizio Silvestri, Amin Mantrach
arXiv Computation and Language
Aug 31

Text Restoration of Ancient Documents with Language Models

The paper explores using language models to restore missing text in damaged ancient manuscripts caused by physical gaps. It tests various scenarios, model architectures, and decoding strategies to handle tokenization mismatches and lacuna length awareness. Results show that while full automation is not yet possible, these tools can effectively aid paleographers, with performance varying by document section and missing text length.

By Shibingfeng Zhang, Edoardo Caraffa, Annafelicia Zuffrano, Maddalena Modesti, Giovanni Colavizza
arXiv Machine Learning
Aug 31

How Far Should Tokenization Go? Predictive Effectiveness and Relational Losslessness

The paper proposes the Effectiveness–Losslessness Framework to guide tokenization in domains beyond language, using predictive codelength as a criterion. It introduces two boundaries: the Fact–Token Boundary, where observable structure should be encoded into tokens, and the Token–State Boundary, where context‑dependent relations should remain for model state rather than being pre‑tokenized. Experiments on symbolic music show that making musical time explicit and applying tonal‑frame canonicalization improve predictive performance, while fixed pitch coordinates and reversible BPE can increase predictive code length, indicating that carrier compaction alone does not guarantee better predictions.

By Yi Wang
arXiv Machine Learning
Aug 31

Post-Training VLMs for Video Mistake Detection

The paper introduces a new protocol, Mistake Detection Video Question Answering (MD‑VQA), to evaluate whether models can determine if a step in a video follows its description, covering both seen and unseen actions. It proposes a post‑training approach for video‑language models that uses a reward function to highlight discrepancies between instructions and video content. Experiments show this method surpasses zero‑shot, fine‑tuned, and other post‑training baselines, especially on unseen procedures, improving performance by up to 11.6% on EP‑VQA.

By Federico Spurio, Olga Zatsarynna, Lars Doorenbos, Emad Bahrami, Gianpiero Francesca, Juergen Gall