Natural language processing

Classical and neural NLP: translation, question answering, tokenization and the evaluation of language understanding.

2,601 stories · RSS feed

arXiv Computation and Language
Sep 1

Reasoning over Grammar: Can Synthetic Linguistic Reasoning Traces Enhance Low-Resource Machine Translation?

The paper explores whether structured linguistic reasoning traces can improve low‑resource machine translation by guiding large language models (LLMs). It proposes a pipeline that automatically generates step‑by‑step reasoning traces from Universal Dependencies treebanks, dictionaries, and grammar‑rule banks, and evaluates these traces in in‑context learning, supervised fine‑tuning, and reinforcement fine‑tuning on Xibe and Chintang. The results show that providing reliable reasoning traces at inference time significantly boosts translation quality, whereas using them as training data yields smaller, less consistent gains, indicating that LLMs can benefit from grammatical guidance but struggle to generate accurate analyses themselves.

By Renhao Pei, Yihong Liu, Sampo Pyysalo, Hinrich Sch\"utze, Shaoxiong Ji
arXiv Computation and Language
Sep 1

QQ: A Language Metadata Toolkit for Multilingual NLP

QQ is a language metadata toolkit designed for multilingual NLP research. It aggregates diverse language metadata into a graph of varieties, scripts, regions, identifiers, names, and relations, and offers access via a Python API, CLI, and browser explorer. The toolkit enables normalization of identifiers, metadata retrieval, relation traversal, and discovery of external resources containing a language, and is demonstrated through audits of the HuggingFace Hub, linking resources with different identifier systems, and generating reproducible language-reporting tables.

By Wessel Poelman, Yiyi Chen, Miryam de Lhoneux
arXiv AI
Sep 1

Do Language Models Reason Across Languages?

The paper investigates whether language models can reason across languages by introducing a two‑hop question answering task that requires inference over two multilingual documents. Results show that models are more sensitive to language variation in answer‑span documents than in bridging documents, and that up to 33% of multilingual cases involve correct final answers despite failing to infer bridging information in the first step. The study also reveals an 18% composition failure rate and proposes a three‑stage SUBQ prompting method that improves accuracy from 10.1% to 66.5%.

By Yan Meng, Wafaa Mohammed, Christof Monz
arXiv Computation and Language
Sep 1

FLAME: A New Dataset on FLemish Accounts of Momentary Experiences

FLAME (FLemish Accounts of Momentary Experiences) is a corpus of nearly 25,000 personal narratives in Belgian-Dutch (Flemish) gathered via experience sampling. The dataset focuses on everyday, culturally grounded themes but presents challenges due to its informal register and low-resource status. Comparative analysis of K‑Means, LDA, and BERTopic shows that BERTopic yields the most coherent and culturally resonant topics according to human evaluation.

By Ratna Kandala, Niels Vanhasbroeck, Katie Hoemann
arXiv Computation and Language
Sep 1

Evidence-Bounded Mental Health Reasoning from Heterogeneous Speech Protocols

The paper introduces Evidence-Bounded Mental Health Reasoning, addressing the problem that current multimodal mental health screening models treat all clinical speech protocols as equally evidential. It presents the Evidence Package Benchmark, comprising 1,870 annotated packages from six diverse protocols, and proposes EviBound, a protocol-aware framework that limits reasoning to valid evidence using a planner, acoustic consensus, and a boundary critic. EviBound outperforms existing omni-modal baselines, achieving a Depression AUROC of 0.8658 with no claim violations.

By Chengyuan Gao, Jiang Wu, Tao Lu, Jiayan Guo, Mingkun Xu, Tianyi Zang, Shangyang Li
arXiv Computation and Language
Sep 1

Hi-Q: Hierarchical Evidence-guided Query Refinement for Multi-Hop Question Answering

Hi-Q is a new framework for multi‑hop question answering that refines queries hierarchically based on evidence retrieved from a corpus. At each node it tests whether the current query unit is supported by evidence; if not, the node is expanded using a dependency‑preserving binary operator and verified for semantic coverage. The resulting query tree grows according to corpus support signals, and Hi‑Q achieves state‑of‑the‑art performance on three multi‑hop QA benchmarks, outperforming both iterative retrieval and graph‑based baselines without constructing a corpus‑wide graph.

By Jueun Kim, Sungho Park, Wook-Shin Han
arXiv AI
Sep 1

Game-Agnostic Value Functions through Automatic JSON Feature Extraction

The paper introduces JSON-Bag VF, a game-agnostic method for training value functions using JSON-Bag prototypes derived from tokenized game trajectories. It demonstrates that Random Forest-based feature selection and game-stage-specific feature selection enhance performance, and that these selections are more critical than prototype-tokenization. Experiments on six tabletop games show that JSON-Bag OSLA outperforms baseline one-step-look-ahead agents in most cases.

By Dien Nguyen, Diego Perez-Liebana
arXiv Computation and Language
Sep 1

Token Counts Are Not Model Lineage: A Frozen-Threshold Holdout Study of Black-Box LLM API Fingerprinting

The study evaluates whether prompt‑token counts can reliably identify the lineage of large language models served via APIs. Using a frozen‑threshold approach on 24 labeled endpoint pairs, the authors find that token‑count consistency perfectly separates development pairs but only half of the holdout pairs meet the strict repeatability criteria, yielding moderate accuracy and perfect specificity. The results confirm token‑count consistency as a fingerprint of shared tokenization stacks but reject it as a standalone test for model‑family attribution.

By Bo Chen
arXiv Computation and Language
Sep 1

Pad\=artha: Ontology-Grounded Fine-Grained NER Benchmark for Classical Sanskrit

Padàrtha is the first ontology‑grounded fine‑grained Named Entity Recognition benchmark for Classical Sanskrit, built on the Mahêbhárata epic. Its tag set, derived from the Nyáya‑Vai’séka ontological system, contains 18 fine‑grained categories under 10 nodes and maps to five standard coarse tags, ensuring compatibility with existing benchmarks. The dataset includes over 12.6K expert‑annotated entries and 108,335 entity mentions across 73,632 verses, plus a 5,000‑verse test set designed to challenge rare mentions, and the study compares generative NER models to traditional architectures, noting performance drops at finer granularity and difficulties with unseen entities.

By Sujoy Sarkar, Pretam Ray, Paramhans Shah, Manoj Balaji Jagadeeshan, Akash Gairola, Arjuna S R, Pawan Goyal
arXiv AI
Sep 1

Responsible Integration of AI in Cancer Genomics: Barriers, Risks, and Pathways to Trustworthy Clinical Translation

arXiv:2608.30912v1 Announce Type: new Abstract: Artificial intelligence (AI) and natural language processing (NLP) are increasingly used to extract, integrate, and interpret biomedical knowledge rele...

By Bahar \.Ilgen, Yiannos Tolias, Denise K\"uhnert, Paraskevi Papadopoulou, Magnus Westerlund, Dominik Heider, Katharina Ladewig, Georges Hattab