Natural language processing

Classical and neural NLP: translation, question answering, tokenization and the evaluation of language understanding.

2,601 stories · RSS feed

arXiv Computation and Language
Sep 14

The Cost of Compression: A Rate-Distortion Limit on Factual Hallucination

The paper presents a rate‑distortion framework for understanding factual hallucination in closed‑book question answering. It shows that even when a fact is observed, limited memory forces it to be stored approximately, leading to errors that can be bounded by a combination of compression distortion and missing coverage. The authors derive a theoretical lower bound on error and validate it with simulations and probes on modern language models.

By Xi Wang, Shijia Xu, Rongfeng Guo
arXiv Computation and Language
Sep 14

CueMem: Cue-Guided Context Reconstruction for Long-Term Conversational Memory

CueMem is a cue‑guided framework for long‑term conversational memory that reconstructs query‑relevant dialogue context from compressed memory records. Instead of treating memory units as self‑contained evidence, it extracts fine‑grained cues linked to their source turns and, at query time, expands from these cues over a turn graph to rebuild a compact evidence context. Experiments on LoCoMo and LongMemEval show that CueMem outperforms baseline memory methods, reduces input tokens and latency, and improves long‑term conversational question answering.

By Changjian Wang, Rongzhen Li, Weili Guan, Shuming Shi, Quan Lu, Ning Jiang
arXiv Computation and Language
Sep 14

EAR: Entity-Aware Partitioning Approach for Retrieval-Augmented Generation Development

The paper introduces EAR, an Entity‑Aware Partitioning approach that improves retrieval‑augmented generation for multiple‑choice question answering by extracting normalized surface anchors from questions, answers, and the corpus. EAR retrieves local windows around matching anchors and can attach a larger parent passage via an extractive summary, reducing retrieved words by 37.5‑40.2% compared to fixed‑size chunks. Experiments on a cleaned MMLU‑style subset with Mistral, Gemma, and DeepSeek show modest accuracy changes, none statistically significant, highlighting EAR’s methodological contribution of compact, inspectable retrieval units.

By Cenab Batu Bora, Oylum Alatl{\i}, Sebnem Bora, Oguz Dikenelli
arXiv Computation and Language
Sep 14

Doc2FRC: Length-Consistent Document-Level Machine Translation via Fixed-Range Chunking

Doc2FRC introduces Fixed-Range Chunking (FRC), a dynamic programming method that partitions documents into chunks of a predefined length interval, ensuring consistent length distributions during training and inference. This approach reduces train-test length mismatch, mitigates n-gram repetition, and improves translation quality for 7B LLMs compared to direct Doc2Doc fine-tuning. Experiments on IWSLT2017 and a new 10-language test set, GlobVDoc, demonstrate that FRC outperforms existing document-level machine translation methods and enhances out-of-distribution translation performance.

By Xiaotian Wang, Youyuan Lin, Zhan Shen, Hitomi Yanaka
arXiv Computation and Language
Sep 14

UrduFactCheck: An Agentic Fact-Checking Framework for Urdu with Evidence Boosting and Benchmarking

The paper introduces UrduFactBench and UrduFactQA, two hand‑annotated benchmarks for claim verification and factual consistency evaluation in Urdu, created through a multi‑stage annotation process with native speakers. It also presents UrduFactCheck, a modular fact‑checking framework that uses both monolingual and translation‑based evidence retrieval to address the scarcity of high‑quality Urdu evidence. Experiments on twelve LLMs show that translation‑augmented pipelines outperform monolingual ones, highlighting ongoing challenges for open‑source models in Urdu.

By Sarfraz Ahmad, Hasan Iqbal, Momina Ahsan, Numaan Naeem, Muhammad Ahsan Riaz Khan, Arham Riaz, Muhammad Arslan Manzoor, Yuxia Wang, Preslav Nakov
arXiv AI
Sep 12

ReactHuman: A Physics-Grounded Benchmark for Human-Like Reactive Decision-Making in Embodied Multimodal LLMs

ReactHuman is a physics‑grounded benchmark that tests whether multimodal large language models (MLLMs) can make immediate, safety‑critical decisions in simulated humanoid scenarios involving sudden household hazards. The benchmark includes 17 event families, over 1,000 reproducible scenes generated from 240 Hz rigid‑body simulation, and a five‑metric suite evaluating reactions on reasonableness, safety, and physical grounding. Evaluation of seven MLLMs reveals that reactive safety remains unsolved, with models frequently mishandling hazards, relying on appearance over motion, and missing key interception points.

By Yizhan Li, Jianxin You, Mengyang Xiong, Yinhuan Chen, Zicheng Zhao, Dekun Wu, Dongqing Zhang, Bang Liu
arXiv AI
Sep 12

ActMap: Single-Pass Uncertainty Quantification from Generation-Time Activation Maps

ActMap is a new white‑box representation that compresses the entire hidden‑state trajectory of a language model during generation into a fixed 12 × 32 × 128 tensor. This compact 96 KiB map can be captured with no overhead and is read by a lightweight Vision Transformer to estimate answer correctness in a fraction of a millisecond. In experiments on short‑answer QA, math, and summarization, ActMap outperforms sampling, token‑probability, attention, and embedding baselines and matches a larger ACT‑ViT detector while achieving lower calibration error on most test pairs.

By Jacopo Dardini (University of Bologna), Roberta Calegari (University of Bologna)
arXiv AI
Sep 12

An AI-Powered Culturally Aware Chatbot for Stress Detection and Wellness Support among Pakistani University Students Using NLP and Machine Learning

The paper presents an AI‑powered, culturally aware chatbot designed to detect stress and provide wellness support for Pakistani university students. Using a Random Forest model trained on 1,100 responses across 20 psychological, physiological, academic, environmental, and social features, the system achieves 89.09% accuracy and a macro F1‑score of 0.89 across three stress severity levels. The model’s outputs are fed to an open‑source large language model via OpenRouter, enabling culturally tailored conversations in English, Urdu, and Roman Urdu, with teacher‑student relationship identified as a key stressor.

By Muhammad Fahad Bashir, Muhammad Afzal
arXiv Computation and Language
Sep 11

Unadapted Multilingual ASR on a Garrusi Kurdish Evaluation Set: A Common-Reference Staged Normalization Analysis

The paper evaluates a multilingual ASR model (MMS‑1B‑all) on a Garrusi Kurdish dataset using a common‑reference staged normalization approach. By normalizing both reference and hypothesis, the authors show that raw Arabic‑script hypotheses yield a 111.70 % WER, which drops to 97.85 % after folding into a reduced orthography, highlighting the impact of orthographic differences on error measurement. A Southern Kurdish fine‑tuned system performs worse, and residual errors are partly due to scoring‑pipeline limitations rather than recognition failures.

By Hiwa Asadpour
arXiv Computer Vision
Sep 11

Towards AI-Driven Policing: Interdisciplinary Knowledge Discovery from Police Body-Worn Camera Footage

The paper introduces an interdisciplinary framework that uses AI and machine learning to analyze police body‑worn camera footage from the Rochester Police Department. It combines image, audio, and natural language processing—including speaker separation, transcription, and large language models—to detect and classify interaction patterns such as respect, disrespect, escalation, and de‑escalation. A custom evaluation pipeline assesses transcription quality and behavior detection accuracy, aiming to support law‑enforcement review, training, and accountability.

By Anita Srbinovska, Angela Srbinovska, Vivek Senthil, Jonathan Bateman, Adrian Martin, John McCluskey, Ernest Fokou\'e
arXiv Computation and Language
Sep 11

Same Day, Same Story; One Day Ahead, a Different Signal: The Dual Validity of Financial Sentiment

The study examines whether financial sentiment tools that are validated against human labels also reliably predict market outcomes. Using a large corpus of securities class action messages linked to abnormal stock returns, the authors compare five sentiment instruments—VADER, Loughran‑McDonald, FinBERT, Twitter‑RoBERTa, and an LLM annotator—within a single pipeline. Results show that the alignment between human agreement and sentiment scores varies with sampling strategy and time horizon: conventional sampling favors same‑day associations, while fixed‑n panels yield similar correlations for both same‑day and one‑day‑ahead predictions, yet overall predictive rankings remain weak.

By AS Aravinthkakshan, Laven Srivastava, Harsh Nandwani
arXiv Machine Learning
Sep 11

Robust Multimodal Sentiment Analysis with Incomplete Modalities via Semantic-aware Completeness based Reconstruction

The paper presents a method for robust multimodal sentiment analysis that handles incomplete or noisy modalities. It introduces a completeness estimation technique to measure how much sentiment-relevant information remains in partial data, guiding the reconstruction of missing semantics. A joint training strategy stabilizes multi-task learning for sentiment prediction and completeness estimation, and experiments on three benchmark datasets show improved semantic reconstruction and sentiment accuracy.

By Han-Jun Choi, Byunggill Joe, Saim Shin, Jin Yea Jang
arXiv Computation and Language
Sep 11

Cross-Lingual Clinical Annotation Projection as Constrained Text Generation: A Six-Language Study

The study investigates whether cross‑lingual clinical annotation projection can be treated as a constrained text‑generation task that preserves the original text while inserting entity tags. Using a workflow that embeds tags directly into immutable target‑language text and then validates them deterministically, the authors evaluated this approach against supervised candidate‑span projection and hybrid ML‑LLM refinement across six languages. Results show that direct LLM projection, particularly with GLM 5.2 and Gemma4:31B, achieves the highest strict F1 scores (up to 0.9201) and outperforms previous methods by 0.0564–0.1512, producing over 55,000 grounded mentions with accurate offsets.

By \'Alvaro Rey-Blanes, Francisco J. Moreno-Barea, Francisco J. Veredas
arXiv Machine Learning
Sep 11

E-CONAN (Entailment, CONtradition And Neutral) Benchmarks: Arabic Textual Entailment and Natural Inference Datasets

E-CONAN introduces Arabic textual entailment and natural inference benchmarks comprising two datasets: E-CONAN-2 (2-way RTE) and E-CONAN-3 (3-way NLI). The datasets are built from automatically-translated pairs, human-validated machine translations, hand-crafted pairs from Arabic teaching books, and rumor-containing news headlines. The authors evaluated nine multilingual pretrained models and five large language models on these benchmarks, demonstrating that E-CONAN offers a more diverse and robust assessment than existing datasets like XNLI and ArNLI.

By Khloud AL Jallad, Nada Ghneim, Ghaida Rebdawi