Natural language processing

Classical and neural NLP: translation, question answering, tokenization and the evaluation of language understanding.

2,564 stories · RSS feed

arXiv Computer Vision
Sep 16

MechReason: Benchmarking Multi-Image Multi-Hop Reasoning in Mechanical Engineering

MechReason is a new benchmark for multi-image, multi-hop reasoning in mechanical engineering, featuring 12,000 question-answer pairs with explicit reasoning chains and 21,000 visual items across nine evidence types. It covers eight task types—explanation, prediction, design, and diagnosis—within four reasoning dimensions, and is constructed through a four-stage pipeline that extracts engineering claims, generates verification-masked questions, and validates multimodal quality. Even state‑of‑the‑art models score only 62.89% on this dataset, highlighting its difficulty.

By Tengyue Wang, Kang An, Chenxu Du, Zhongyu Yang, Yuanchi Zhu, Xinqi Yang, Hebao Zhu, Ziliang Wang, FaQiang Qian, Yunli Yang, Qibing Ren
arXiv AI
Sep 16

Diagnosing the Fact-Grounding Gap in Multi-Hop Question Answering

The paper investigates multi‑hop question answering systems and identifies two distinct failure modes: retrieval failures, where the necessary passage is not retrieved, and extraction failures, where the passage is retrieved but the required fact cannot be extracted—a phenomenon termed the fact‑grounding gap. Across three standard benchmarks, extraction failures account for nearly half of all per‑hop deficiencies and are invisible to standard retrieval metrics, remaining unresolved by retrieval‑only interventions. The study shows that these two bottlenecks require different solutions, a distinction currently missing from evaluation practices.

By Kevin Mo, Nathan Mo, Richard Zhu
arXiv AI
Sep 16

Vision And Text Transformer For Predicting Answerability On Visual Question Answering

The paper introduces VT-Transformer, a model that predicts answerability scores for Visual Question Answering by treating the task as a regression problem rather than a binary classification. It leverages visual and textual features within a Transformer architecture and demonstrates improved performance and robustness on the VizWiz 2020 dataset compared to existing baselines.

By Tung Le, Huy Tien Nguyen, Le Minh Nguyen
arXiv Computation and Language
Sep 16

Nameless Tokenization: A Lossless Tokenizer-Level Defense Against Control-Token Forgery in Open-Weight LLMs

The paper examines how open‑weight language models expose the control tokens used in chat templates, allowing attackers to forge turn boundaries that the model treats as legitimate. An audit of 256 deployed tokenizers shows all are vulnerable, and the commonly recommended flag fails to protect 56.6% of cases. The authors introduce nameless tokenization, which removes surface strings for control identifiers while preserving their internal representation, achieving identical token streams on clean data and significantly improving accuracy on delimiter‑bearing text.

By Kisu Yang, Yoonna Jang, Heuiseok Lim
arXiv Machine Learning
Sep 16

Single Document Extractive Summarization using Domination in Hypergraph

The paper proposes a new approach to single-document extractive summarization by constructing a sentence hypergraph where sentences are nodes and keywords or named entities are hyperedges. A greedy algorithm is then used to find a dominating set of this hypergraph, which yields the sentences that compose the summary. The study compares this hypergraph-based method with existing graph-based summarization techniques.

By Aamir Miyajiwala, Aabha Pingle, Sheetal Sonawane, Surajit Kr. Nath
arXiv Machine Learning
Sep 16

AsyncCouple-Flow: Asynchronous Cross-Modal Coupling and Flow Matching for Spatio-Temporal Forecasting

AsyncCouple-Flow introduces a new framework for multi‑modal spatio‑temporal forecasting that tackles three key challenges: differing sampling rates, missing modalities, and autoregressive error accumulation. It employs a Modality‑Aware Token Sparsification module to produce equal‑length sequences, an Asynchronous Cross‑Modal Coupling Graph to fuse data under arbitrary asynchrony and missingness, and a Flow‑Matching Forecasting Head that models multi‑step prediction as a conditional ODE. Experiments on weather and traffic datasets demonstrate that the method outperforms state‑of‑the‑art baselines and remains robust even when up to two modalities are missing.

By Zhixiang Wu, Yining Liu, Bo Zhao, Szu-Yu Chen, Huiran Duan, Chu Lin, Chuanguang Yang
arXiv Computer Vision
Sep 16

Not Another Text Benchmark: Putting the "Visual" Back in Visual Question Answering for Large Video Models

The paper introduces three new vision‑centric evaluation benchmarks—temporal frame retrieval, video future prediction, and causal memory distortion—to assess visual question answering in large video models. Unlike traditional benchmarks that rely on text-based multiple choice questions, these tasks require models to reason directly from visual inputs. The authors find that current state‑of‑the‑art models struggle with visual queries, highlighting a gap in visual understanding that future research should address.

By Rwiddhi Chakraborty (Oliver), Yinong (Oliver), Wang, Cheng Zhang, Fan Bai, Zhuoran You, Michael Kampffmeyer, Yong Jae Lee, Fernando De la Torre, Robert Jenssen
arXiv Machine Learning
Sep 16

Differentially Private Semantic Plans for Aggregate Insight Generation

The paper introduces ‘DP-SPIN’, a trusted‑curator framework that generates differentially private semantic plans for aggregate insight generation. ‘DP-SPIN’ maps each record to a bounded sparse nonnegative vector over pre‑defined semantic concepts, sums these vectors into a semantic sketch, and releases a noisy plan containing admitted concepts and their masses. The framework provides user‑level privacy by clipping each user’s contribution and ensures that the final summary is differentially private through post‑processing, with guarantees established under both add/drop and replacement adjacency.

By Behrooz Razeghi
arXiv Machine Learning
Sep 16

HUMAID-NER: A Disaster Tweet Dataset for Joint Named Entity Recognition and Event Classification via Uncertainty-Weighted Multitask Learning

HUMAID-NER is the first named entity recognition dataset built on the HumAID benchmark, comprising 60,000 English disaster tweets with approximately 175,000 labeled entity spans across ten operationally motivated entity types. The dataset was created using a reproducible three‑stage hybrid pipeline that combines a spaCy transformer model, disaster‑domain EntityRuler patterns, and structured regular expressions with priority‑based overlap resolution. A joint multitask learning framework using a shared RoBERTa‑large encoder and homoscedastic uncertainty weighting achieves an NER span micro‑F1 of 0.841 and classification macro‑F1 of 0.761, and the authors provide a real‑time web dashboard, dataset, models, and pipeline code for reproducibility.

By Aijaz Ali, Nazish Basir, Sarfaraz Nawaz, Danish Nazir Arain, Haris Ali
arXiv Computation and Language
Sep 16

ReMova: Fine-tuning LLMs for English to Belarusian translation

The paper introduces ReMova, a pipeline for cleaning Belarusian data and fine‑tuning large language models (LLMs) for English‑to‑Belarusian translation. It uses a correction tool to handle the two orthographies of Belarusian, remove noise, filter out interference from other languages, and correct common misspellings found online. Ablation experiments on unfiltered data show that filtering benefits all fine‑tuned models, with LLM‑based models gaining about twice as much as a dedicated encoder‑decoder MT system, highlighting data quality as a key bottleneck for Belarusian MT.

By Mikita Pilinka, Aliaksandr Kliuje\u{u}, David Samuel, Yves Scherrer