Natural language processing

Classical and neural NLP: translation, question answering, tokenization and the evaluation of language understanding.

2,564 stories · RSS feed

arXiv AI
Sep 28

Strategic Self-Consistency

The paper "Strategic Self-Consistency" investigates how large language model providers might exploit the self‑consistency technique—generating multiple reasoning paths and selecting the majority answer—to overcharge users. The authors present a simple, efficient algorithm that strategically generates and reorders extra reasoning paths so that each appears necessary for the majority vote, thereby evading detection by auditors. Experiments on Llama, Qwen, and DeepSeek-R1 models across math, science, and QA benchmarks show that the added paths follow a heavy‑tailed distribution and that significant overcharging can persist even under stringent audits with a false‑positive rate below 0.1.

By Tori Qiu, Ander Artola Velasco, Manuel Gomez-Rodriguez
arXiv AI
Sep 28

PolicyAttention: Softmax Attention Implements Policy Mirror Descent for Closed-Loop Control

The paper investigates whether causal softmax attention can realize policy mirror descent (PMD) as a repeated controller rather than a one‑step algebraic identity. It constructs a fixed causal‑softmax actor–environment–one‑step‑critic protocol, detailing actor, routing, sampling, and normalization residuals, and shows that a frozen one‑step audit model closely approximates PMD. Empirical results demonstrate that the learned actor with an exact one‑step critic achieves median policy loss only about 5% higher than the exact PMD oracle across multiple control settings.

By Yuhe Sui, Yingzhi Tang, Shufang Chen
arXiv AI
Sep 28

TrafficImag: A Benchmark for Counterfactual Roadside Traffic Video Generation

TrafficImag is the first benchmark designed to evaluate counterfactual roadside traffic video generation, combining a large roadside dataset with an executable protocol that supports behavior reasoning, intervention-aware image editing, and conditional video generation. Each intervention is encoded as an actor-level program specifying target actor, intended behavior, legal route, interaction order, and temporal constraints, allowing a unified evaluation across diverse foundation models. The benchmark assesses four validity dimensions—initial-state correctness, route and behavior validity, interaction consistency, and non-target preservation—and reports that the best models achieve 80.4% macro F1 for reasoning and 55.0% end-to-end success when using a complete condition interface.

By Xiangyu Li, Tianyi Wang, Zhihao Dou, Christian Claudel, Zhaomiao Guo
arXiv AI
Sep 28

Acoustic-to-Text KV Compression for Full-Duplex Speech Models

The paper introduces acoustic-to-text KV compression for full‑duplex speech models, converting acoustic key‑value states into compact textual memory during listening‑time slack. When the KV cache exceeds a target budget, older acoustic states are evicted while transcripts and recent acoustic context are retained. Experiments on ten‑minute LongSpeech sessions show a 64.6% reduction in peak streaming KV‑cache size and improved transcription, temporal question answering, and summarization, with comparable pause‑handling, turn‑taking, and interruption performance in Full‑Duplex‑Bench.

By Yejin Lee, Seungbeom Kim, Yongha Lee, Kyuhong Shim
arXiv AI
Sep 28

FlyAOC: Evaluating Agentic Ontology Curation of Drosophila Scientific Knowledge Bases

FlyAOC is a benchmark that tests AI agents on end‑to‑end ontology curation of Drosophila scientific literature. Given a gene symbol, a brief description, a large paper corpus, and ontology resources, agents must search for evidence and produce structured annotations such as function terms, expression patterns, and historical synonyms. The benchmark contains 7,397 expert‑curated annotations across 100 genes and evaluates different agent harnesses, revealing system‑level failure modes that single‑task evaluations miss.

By Xingjian Zhang, Sophia Moylan, Ziyang Xiong, Qiaozhu Mei, Yichen Luo, Jiaqi W. Ma
arXiv AI
Sep 28

AcuityBench: Evaluating Clinical Acuity Identification and Uncertainty Alignment

AcuityBench is a new benchmark that tests whether language models can correctly identify the urgency of medical care needed from user presentations. It unifies five public datasets—user conversations, online forum posts, clinical vignettes, and patient portal messages—under a shared four-level acuity framework, providing 914 cases for evaluation. The benchmark supports both explicit four-way classification and free-form conversational responses, revealing that models vary widely in accuracy and that conversational formats reduce over-triage but increase under-triage, especially for high-acuity cases.

By Robin Linzmayer (Department of Computer Science, Columbia University, Department of Biomedical Informatics, Columbia University), Georgianna Lin (Department of Biomedical Informatics, Columbia University), Di Coneybeare (Department of Emergency Medicine, Columbia University Irving Medical Center), Jason Chu (Department of Emergency Medicine, Columbia University Irving Medical Center), Trudi Cloyd (Department of Emergency Medicine, Columbia University Irving Medical Center), Manish Garg (Department of Emergency Medicine, Columbia University Irving Medical Center), Miles Gordon (Department of Emergency Medicine, Columbia University Irving Medical Center), Elizabeth Hartofilis (Department of Emergency Medicine, Columbia University Irving Medical Center), Benjamin Hong (Department of Emergency Medicine, Columbia University Irving Medical Center), Ashraf Hussain (Department of Emergency Medicine, Columbia University Irving Medical Center), Eugene Y. Kim (Department of Emergency Medicine, Columbia University Irving Medical Center), Oluchi Iheagwara King (Department of Emergency Medicine, Columbia University Irving Medical Center), Ross McCormack (Department of Emergency Medicine, Columbia University Irving Medical Center), Erica Olsen (Department of Emergency Medicine, Columbia University Irving Medical Center), John K. Riggins Jr (Department of Emergency Medicine, Columbia University Irving Medical Center), Mustafa N. Rasheed (Department of Emergency Medicine, Columbia University Irving Medical Center), Dana L. Sacco (Department of Emergency Medicine, Columbia University Irving Medical Center), Vinay Saggar (Department of Emergency Medicine, Columbia University Irving Medical Center), Osman R. Sayan (Department of Emergency Medicine, Columbia University Irving Medical Center), Amit Shembekar (Department of Emergency Medicine, Columbia University Irving Medical Center), Janice Shin-Kim (Department of Emergency Medicine, Columbia University Irving Medical Center), Wendy W. Sun (Department of Emergency Medicine, Columbia University Irving Medical Center), Bernard P. Chang (Department of Emergency Medicine, Columbia University Irving Medical Center), David Kessler (Department of Emergency Medicine, Columbia University Irving Medical Center), No\'emie Elhadad (Department of Computer Science, Columbia University, Department of Biomedical Informatics, Columbia University)
arXiv AI
Sep 28

NaijaNLP: A Survey of Nigerian Low-Resource Languages

The paper surveys NLP research on Nigeria’s three major low‑resource languages—Hausa, Yoruba, and Igbo—covering over 500 languages spoken by 175 million people. It reviews 293 studies, finding that only 27.6% produced new linguistic resources, indicating a heavy reliance on repurposing existing data. The authors highlight under‑explored challenges such as morphological analysis and diacritic representation, and call for collaborative resource enrichment and community support to advance NaijaNLP and low‑resource NLP more broadly.

By Isa Inuwa-Dutse
arXiv AI
Sep 28

Geometry-Aware Hyperbolic Residual-Quantized Variational Autoencoders

The paper introduces a geometry-aware hyperbolic residual quantization method for variational autoencoders, addressing inconsistencies that arise when extending residual vector quantization to hyperbolic space. It restores telescoping behavior in the forward pass via Hyperbolic Residual Aggregation and improves gradient flow in the backward pass with a discounted Hyperbolic Straight-Through Estimator. Experiments on hierarchical prediction, recommendation, image tokenization, and neural audio coding demonstrate enhanced stability and structural organization of hyperbolic residual codes, while highlighting a trade‑off between compression efficiency and hierarchical organization.

By Alessio Colombo, Melika Ayoughi
arXiv Computation and Language
Sep 28

KuaFu: Compressing Long User Behavior into Understanding at Billion Scale

KuaFu is a unified behavior‑compression layer that reduces each user behavior item to 2–4 tokens, dramatically shrinking per‑item cache size while preserving fidelity through a four‑stage training process. In production across four profiling tasks, it matches or outperforms uncompressed single‑task models, boosts GPU throughput by 37–350%, and saves 190 GPUs. On public benchmarks it consistently beats prior compressors at the same compression ratio, and on RecBench a 4B KuaFu model outperforms its 8B counterpart by 1.90 points, contributing to a 1.37% lift in overall GMV on Tencent’s advertising and recommendation platform.

By Jiahao Hui, Lin Zhu, Yishen Hu, Jingdong Shu, Zetai Jiang, Xining Ran, Ben Tan, Yeshou Cai, Gong Chen, Haijie Gu, Jie Jiang
arXiv Computation and Language
Sep 28

Prompt Injection Detection for Email Agents Through Attack Chain Modeling

The paper introduces a prompt‑injection detection framework for email assistants that models attacks as a chain of stages. It combines a text detector, stage‑specific verifiers, rule‑based risk signals, user intent consistency checks, and a logistic decision policy. Experiments on five benchmarks show the framework outperforms pretrained detectors, achieving a mean F1 of 0.406 versus 0.216, and demonstrate that training on benign emails resembling attacks reduces false alarms.

By Ahmad Hashmi, Dhyey Patel, Yunting Yin
arXiv Computation and Language
Sep 28

Asking For An Old Friend: Diagnosing and Mitigating Temporal Failure Modes in LLM-based Statutory Question Answering

Large language models (LLMs) are increasingly used for legal research, but their fixed training cutoffs and reliance on static knowledge clash with the evolving nature of statutory law. This study introduces a benchmark of 312 expert‑validated, time‑sensitive German statutory QA pairs that examine two temporal failure modes: post‑cutoff staleness and recency bias. Five LLMs were evaluated under four inference settings, and the results show that retrieval‑augmented approaches that enforce temporal validity significantly improve performance, while web search yields unstable gains and a pronounced recency bias.

By Max Prior, Andreas Schultz, Matthias Grabmair
arXiv Computation and Language
Sep 25

SemMSA: Latent Semantic-Aided Robust Multimodal Sentiment Analysis with Incomplete Data

SemMSA introduces a latent semantic‑aided framework for multimodal sentiment analysis that leverages large language models to generate rich sentiment‑relevant semantics. The method employs Cross‑modal Semantic Refinement (CSR) to fuse visual, acoustic, and language features in a frozen LLM embedding space, and Cross‑modal Spectral Alignment (CSA) to align these refined semantics with all modalities via spectral enhancement of kernel Gram matrices. Experiments on SIMS, MOSI, and MOSEI benchmarks show that SemMSA achieves state‑of‑the‑art performance.

By Wenhao Li, Zhibin Wu, Chong Xiao, Qiangchang Wang
arXiv Computer Vision
Sep 25

CinematicVQA: Benchmarking Film-Grammar Reasoning in Large Vision-Language Models

CinematicVQA is a new benchmark for evaluating large vision‑language models on film‑grammar reasoning. It introduces the Cinematic Scene Graph, a structured representation linking filming techniques to perceptual effects and narrative functions, and tests models on tasks beyond low‑level technique recognition. The study finds a semantic gap where models excel at describing visuals but struggle to identify underlying techniques, and shows that fine‑tuning improves performance on narrative function and multi‑hop reasoning.

By Shuo Xing, Pooja Verlani, Balu Adsumilli, Zhengzhong Tu
arXiv Computer Vision
Sep 25

Direction-Scale Decomposition in Action Representation: Rethinking What to Tokenize for Vision-Language-Action Models

The paper introduces Direction-Scale Decomposition (DSD), an action representation that separates translation and rotation increments into direction and scale components before tokenization. DSD is evaluated with uniform binning and a B-spline tokenizer (BEAST) in both simulation and real-world manipulation tasks, showing improved success rates on LIBERO and SimplerEnv, especially under mixed-dataset training. Real-robot experiments confirm performance gains with and without robotics pretraining, supporting DSD as an effective representation for discrete-token vision-language-action models.

By Yufei Duan, Hang Yin, Alberta Longhini, Chao Tang, Danica Kragic
arXiv Computer Vision
Sep 25

GeoNLI - A Natural Language Interpreter for Satellite Imagery

GeoNLI introduces a unified, modular pipeline that combines advanced SAM variants with multimodal large language models to perform satellite image captioning, visual question answering (VQA), and visual grounding. The EarthMind model achieves strong results on captioning and VQA, while multiple RemoteSAM-SAM and DiffuSAM pipelines are used for grounding, ultimately employing a majority‑voting ensemble across several models. The system reports 82% captioning accuracy, 83.32% VQA accuracy, and 64.94% grounding accuracy, demonstrating improved consistency over task‑specific approaches.

By Ashutosh Gandhe, Anupam Rawat, Geet Sethi, Kabir Nasiruddin, Madhav Kotecha, Panav Shah, Rakshit Sawarn, Soumitra Nayak
arXiv Computer Vision
Sep 25

PROVE: Proof-guided Regime-aware Operator Verification for Hallucination Detection in Medical Visual Question Answering

PROVE is a black‑box hallucination detector for medical visual question answering that tailors its verification strategy to each question’s evidential structure. It classifies questions into three regimes, activates a subset of five operators per regime, and calibrates operator importance using deterministic question‑answer features to produce a risk score. On 8048 test samples across three medical VQA benchmarks and four state‑of‑the‑art vision‑language models, PROVE achieves an AUROC of 0.821, surpassing the best baseline by 0.159 with consistent improvements across all models and datasets.

By Keyang Zhou, Siyi Li, Zhongnan Shi, Qichao Ying, Wei Tang, Zhenxing Qian
arXiv AI
Sep 25

The Fellowship of the Query: Learning Retrieval Actions

The paper investigates how trajectory fine‑tuning can enhance small language models (SLMs) as next‑action controllers in retrieval‑augmented question answering. By building a seven‑way action‑prediction task from teacher search traces, the authors fine‑tune SLMs and cross‑lingual SLMs (xSLMs) using LoRA and evaluate on 1,646 held‑out examples, achieving a macro‑F1 of 0.6536 with Granite 4.1 3B. In an end‑to‑end controller/generator swap experiment on 149 trajectories, the fine‑tuned model improves Exact Match from 0.7530 to 0.7946 and token F1 from 0.7783 to 0.8295, demonstrating that trajectory supervision boosts action prediction and evidence‑recording behavior.

By Mohammed Al-Maamari, Saber Zerhoudi, Michael Granitzer, Jelena Mitrovi\'c
arXiv AI
Sep 25

Reinforcement Learning with Verifiable Rewards for Small Search Agents

The paper introduces Reinforcement Learning with Verifiable Rewards (RLVR) applied to small search agents, specifically training a Qwen3.5-0.8B model with Group Relative Policy Optimization and an interleaved Wikipedia-search tool on the MuSiQue dataset. Experiments varying reward shapes across three seeds show that RLVR can achieve a 3.8‑fold improvement over an untrained baseline, with the best run reaching a 0.352 average exact match. The study finds that the sparse exact‑match reward, standard in larger models, performs poorly for small models, indicating that reward design must be tailored rather than scaled down from large‑model recipes.

By Gaurisankar Jayadas, Aske Plaat, \'Alvaro Serra-G\'omez, Sandheep P
arXiv AI
Sep 25

From Text Decisions to Pixels: An Study of Jev-Style Visual Choice Model

The paper introduces PixelJev, a native-image decision interface that combines an image, a task instruction, and a runtime candidate set into a structured choice and candidate-conditioned probabilities using small open multimodal models. It unifies recognition and multiple-choice visual question answering via a language-model readout, offering options for frozen inference, language-side adaptation, and held-out calibration. Across seven benchmarks, 64-shot source adaptation significantly boosts Pets accuracy from 60.13% to 92.40%, and the system supports both VQA tasks with frozen inference, while also highlighting areas for improvement such as schema robustness and cross-family transfer.

By Xunlan Zhou, Xianliang Yang, Li Zhao
arXiv Computation and Language
Sep 25

EnSiTa - A Trilingual Multi-Domain Parallel Dataset and Benchmark for Domain-Specific Machine Translation

EnSiTa is a trilingual multi‑domain parallel dataset and benchmark for English, Sinhala, and Tamil. It contains human post‑edited training data across seven domains and professionally translated test sets for those domains plus an additional one, all produced through a multi‑year, rigorously quality‑controlled process. The authors use EnSiTa to conduct a comprehensive study of domain‑specific machine translation across six language directions, comparing from‑scratch Transformers, pre‑trained models, and decoder‑only LLMs under various training‑data sizes, model scales, and domain settings.

By Surangika Ranathunga, Nisansa de Silva, Aloka Fernando, Kavindu Warnakulasuriya, Isuru Wijesiri, Menan Velayuthan, Charitha Rathnayaka, Thivaharan Varatharajan, Sajeevi Silva, Piumi Kandanaarachchi, Uthayasanker Thayasivam