Natural language processing

Classical and neural NLP: translation, question answering, tokenization and the evaluation of language understanding.

2,641 stories · RSS feed

arXiv Computation and Language
Aug 24

Trust Stack for Mental Health AI: A Survey of Calibration across Human, Interaction, and AI Layers

The paper surveys 61 studies on mental‑health AI and identifies a misalignment in how trust is evaluated across disciplines. It proposes a three‑layer framework—human‑oriented, interaction‑oriented, and AI‑oriented trust—and maps stakeholder perspectives onto these layers. The authors argue that future research should focus on calibrating human trust to actual interaction and AI trustworthiness rather than merely maximizing perceived trust.

By Xin Sun, Yue Su, Yifan Mo, Qingyu Meng, Yuxuan Li, Min Chen, Mengyuan Zhang, Saku Sugawara, Charlotte Gerritsen, Sander L. Koole, Koen Hindriks, Jiahuan Pei
arXiv AI
Aug 24

Trilingual Topic Modeling of Sri Lankan Parliamentary Debates

The paper presents an end‑to‑end framework for extracting and clustering trilingual Sri Lankan parliamentary debates in Sinhala, Tamil, and English. Using LLM‑based text extraction, multilingual embeddings, and density‑based clustering, the authors recover 30 macro‑topics with a cluster purity of 0.673. The temporal patterns of these topics align with major national events such as the 2019 Easter attacks and the 2022 economic crisis, demonstrating the method’s effectiveness where traditional LDA fails.

By Himath Dhanapala, Haren Daishika, Himandhi Kuruppu, Sithija Seneviratne, Ashini Kavindya, Patalee Narasinghe, Sandeepa Weerasekara, Nisansa de Silva, Sandareka Wickramanayake
arXiv AI
Aug 24

PSK at WMT 2026 MIST: Task-Specialized QLoRA Adapters for Multilingual Summarization and Question Answering

The PSK submission to the WMT 2026 Multilingual Instruction Shared Task employs a 3.35B‑parameter Tiny Aya Global model enhanced with three QLoRA adapters, each dedicated to a specific task: multilingual summarization, passage‑based question answering, and filtered standalone question answering. The summarization adapter is trained on multilingual document‑summary pairs, including scientific papers with author‑written abstracts, and outperforms a multitask adapter trained solely on organizer data on a held‑out split. For open question answering, results vary with answer length and evaluation method, prompting the submission of three systems that share the same context and summarization adapters but differ in their open‑QA adapters.

By Srikar Kashyap Pulipaka
arXiv AI
Aug 24

SAGE: A Unified Algebra and Self-Adaptive Execution for AI Functions in SQL

SAGE (Self-Adaptive Generative Execution) introduces a unified framework for integrating AI functions into SQL by defining three typed primitives—AI_SCALAR, AI_AGG, and AI_JOIN—that correspond to the relational roles of transforming rows, aggregating groups, and joining row pairs. The framework standardizes a confidence-gated execution interface and tailors physical strategies to each primitive’s shape, with AI_JOIN employing predicate analysis and a recipe card to select optimal execution plans. Evaluations across scalar, aggregate, and join workloads demonstrate that SAGE consistently improves execution quality and efficiency, achieving the best overall SemBench performance and dramatically reducing model calls in factorable joins.

By Xiangqi Wang, Nhan H. Pham, Oktie Hassanzadeh, Dharmashankar Subramanian, Xiangliang Zhang
arXiv AI
Aug 24

Dual-Cache Latent Space Communication between Heterogeneous Language Models

The paper introduces XKV, a latent protocol that enables efficient communication between heterogeneous language models by translating a sharer's key‑value cache into a receiver's context. XKV overcomes limitations of prior methods by jointly pooling both caches, reconciling differing layer depths, and allowing each receiver position to retrieve its own residual in native KV geometry. Across 45 dataset‑model pairings, XKV outperforms previous protocols and text communication while using fewer parameters and achieving faster translation times.

By Jiyao Liu, Qi Zhang, Yaoyi Jia, Ziwen Kan, Song Wang
arXiv Machine Learning
Aug 24

Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs

The paper introduces Llama-Mobile, a framework that quantizes vision‑language models for efficient mobile deployment. It uses a quantization pipeline that generates training data from the model itself, eliminating the need for the original training setup, and employs a novel 2.7‑bit‑per‑parameter format optimized for Arm CPUs. Applying this method, the authors compress the Llama 3.2 11B Vision Instruct model to 3.7 GB with 8‑bit activations while maintaining strong performance on visual question answering tasks.

By Luka Ribar, Jeevan Bhoot, Douglas Orr
arXiv Computation and Language
Aug 24

Granuscore: A Reference-Free Measure of Granularity for Text Analysis and Question Answering

Granuscore is a reference‑free metric that measures the granularity of text by exploiting the structure of a hierarchical embedding space. It successfully reproduces known hierarchical orderings on the Granola‑EQ dataset, distinguishes granularity across different discourse contexts, and explains sentence‑specificity variations beyond sentence length. The authors also apply Granuscore to four question‑answering benchmarks, revealing systematic differences in granularity among questions, gold answers, and model outputs, thereby offering a new lens for assessing QA dataset difficulty.

By Lukas Ellinger, Alexander Fichtl, Miriam Ansch\"utz, Georg Groh
arXiv AI
Aug 24

The Divergence Hypothesis: Unmasking Lexical Interference and Label Bias in Mental Health NLP

The paper introduces TSS (Triple-Stream Stress probe), a diagnostic framework that splits text into lexical, morpho-syntactic, and psycholinguistic style channels to analyze mental health NLP classifiers. Across four English datasets, TSS uncovers a lexical interference effect where adding lexical features harms performance on human-labeled data but not on auto-labeled data, and proposes the Degree of Divergence (DoD) statistic to audit label-source bias. The study demonstrates that style features largely remain effective even after masking content words, emphasizing that shortcut learning is label-source specific rather than clinically relevant.

By Moustafa Yehia Hassan
arXiv AI
Aug 24

StateSight: Benchmarking Latent Spatial-State Reconstruction in Vision-Language Models

StateSight is a new benchmark designed to isolate and evaluate the ability of vision‑language models to reconstruct latent spatial structure from a single image. It consists of three procedurally generated task families—cube‑net opposite‑face reasoning, occluded cube‑tower counting, and 4‑neighbor connected‑component counting—each with 300 deterministic prompts and exact‑match scoring. The benchmark also includes a companion dataset, StateSight‑Steps, with 900 image‑text examples and 3,600 intermediate visual states to aid analysis of reconstruction errors.

By Michelle Lin
arXiv AI
Aug 24

SCOPE: A Generative Approach for LLM Prompt Compression

SCOPE is a training‑free generative prompt‑compression framework that reduces LLM input length by chunking a prompt into semantically coherent segments, rewriting each chunk to be more concise, and then reconstructing a coherent prompt. Unlike token‑removal methods, SCOPE’s chunk‑level rewriting preserves critical information and text coherence, and includes optimization techniques for finer‑grained control of compression ratios. Extensive evaluations on question‑answering and summarization tasks show that SCOPE consistently outperforms selective compression baselines, especially at high compression ratios.

By Tinghui Zhang, Yifan Wang, Daisy Zhe Wang
arXiv AI
Aug 24

STAR-OPD: Structured Aspect-Cascade-Aware On-Policy Reward Distillation for ABSA Quadruple Extraction

The paper introduces STAR-OPD, a structured aspect‑cascade‑aware on‑policy reward distillation method for aspect‑based sentiment analysis (ABSA) quadruple extraction. It addresses a specific failure mode where distilled models produce structurally invalid target‑aspect bindings, leading to hallucinated targets and corrupted downstream predictions. By training on student rollouts and applying set‑structured rewards that enforce binding consistency, target grounding, and fine‑grained aspect disambiguation, STAR‑OPD outperforms both off‑policy and generic on‑policy baselines on E‑ABSA20K and SemEval‑2014, reducing target hallucination and improving performance on structurally hard cases.

By Tong Sun, Mingyang Ma, Jiayang Yu
arXiv Machine Learning
Aug 21

LODESTAR: Robust Entropy-Based Answer Selection in Retrieval-Augmented Generation for Question Answering -- Directing Frozen-LLM Entropy with a Reinforcement-Learned Prompt Polarizer under Misleading Passages

arXiv:2608. 11922v2 Announce Type: replace-cross Abstract: Predictive-distribution entropy is a strong answer-selection rule in retrieval-augmented generation (RAG) for question answering: across five QA benchmarks, selecting the answer a frozen respondent LLM produces with the lowest answer-token entropy lifts mean $F_1$ from 0.

By Hung-Chun Hsu, Po-Jen Ko, Che-Cheng Wu, Li-Yang Chang, Chuan-Ju Wang