Natural language processing

Classical and neural NLP: translation, question answering, tokenization and the evaluation of language understanding.

2,564 stories · RSS feed

arXiv Computer Vision
Sep 17

Sim-to-Real Traffic Scene Understanding by Decoupling Semantics from Caption Generation with V-JEPA

The paper presents a decoupled framework for sim-to-real traffic scene understanding, separating semantic fact extraction from caption generation. It uses a frozen V-JEPA encoder for predictive scene representations and a lightweight Llama-based predictor for VQA, followed by a training-free structured refinement that leverages statistical priors, inter-question relationships, and temporal consistency. The refined facts are then fed to Qwen3-VL-8B to produce pedestrian and vehicle descriptions, achieving top performance on the 2026 AI City Challenge Track 2 benchmark with 87.09% VQA accuracy and an overall S2 score of 60.0853.

By Nguyen Hoai Thuong Bui, Thanh Nguyen Vo, Trinh Tra Giang Nguyen, Ha Duc Bui
arXiv AI
Sep 17

Version- and Scope-Aware Question Answering over Normative Documents: A Deployed System and an End-to-End Evaluation at Production Scale

The paper presents a deployed system for answering questions over normative documents that is aware of document version and scope. It evaluates a hosted retrieval service against a governed system that applies explicit rules for version and scope resolution, finding the governed system achieves a higher score (97.7 vs 88.1). The study includes a public benchmark, evaluation scripts, and reports commercial deployment metrics, such as 1,126 users and 100,000 calls per day by April 2026.

By Liuyin Wang, Shuaipeng Jin, Jiwei Shi, Jensen Hsu
arXiv Computation and Language
Sep 17

Divide and Conquer: Mixture-of-Bottleneck Experts in Informative Ordinal Space for Video-based Multimodal Sentiment Analysis

The paper proposes a Mixture-of-Bottleneck (MoB) framework for video-based multimodal sentiment analysis that treats sentiment as an ordinal regression problem, splitting it into polarity recognition and intensity prediction. MoB assigns modality‑specific latent experts to each sub‑task, learns compact, task‑relevant representations via an information bottleneck, and fuses these experts with a multimodal bottleneck routing module and hard mining strategy. Experiments on four datasets and language models demonstrate that MoB captures fine‑grained intra‑ and inter‑modal dynamics, improving performance and enabling more trustworthy localization of nuanced sentiment signals.

By Ronghao Lin, Qiaolin He, Zefeng Lu, Yichu Liu, Li Huang, Sijie Mai, Haifeng Hu, Yap-peng Tan
arXiv Computation and Language
Sep 17

Beyond frequency measures: Can contextual embeddings capture meaning change in scientific texts?

The study investigates whether contextual embeddings can detect meaning changes in scientific terminology beyond traditional frequency counts. Using Astrophysics and NLP corpora from 2010 to 2024, the authors extract candidate terms with KeyBERT, filter for significant frequency rises, and then evaluate semantic drift via multiple embedding‑based metrics. Results show that frequency methods slightly outperform embedding metrics in aligning with expert judgments, yet embedding‑only detections (e.g., "primordial black holes") reveal critical conceptual shifts missed by frequency alone, suggesting complementary value.

By Jianying Liu (STL, BETA, CEIPI), Kim Gerdes (LISN, Qatent, STL), Jean-Marc Deltorn (CEIPI)
arXiv Machine Learning
Sep 17

Instrument Classification of Solo Sheet Music Images

This paper investigates instrument classification using solo sheet music images rather than audio. It converts images into sequences of musical words via bootleg score representation and treats the task as text classification, training AWD‑LSTM, GPT‑2, and RoBERTa models on IMSLP data for eight instruments. Pretraining on unlabeled data and fine‑tuning improves RoBERTa’s accuracy from 34.5% to 42.9%, and two proposed data‑augmentation methods raise accuracy by an additional 15%.

By Kevin Ji, Daniel Yang, TJ Tsai
arXiv Machine Learning
Sep 17

Patch the Distribution Mismatch: RL Rewriting Agent for Stable Off-Policy SFT

The paper introduces a reinforcement‑learning based rewriting agent that rewrites training data to reduce the distribution mismatch between supervised fine‑tuning (SFT) and a language model’s generation distribution. By formulating data rewriting as a policy‑learning problem, the authors train a lightweight LoRA rewriting policy that optimizes alignment with question‑answering style, maintains semantic diversity, and enforces task consistency. Experiments on three instruction‑tuned backbones show that models fine‑tuned with the rewritten data achieve downstream performance comparable to standard SFT while mitigating degradation on non‑downstream benchmarks, and preliminary tests suggest the policy can transfer across domains such as logical reasoning and medical question answering.

By Jiacheng Wang, Zhijie Liu, Ping Jian, Zirong Chen, Ke Ren Liao, Zhen Yang, Zhongbin Guo
arXiv Computation and Language
Sep 17

A Systematic Review of NLP for Ghanaian Languages: Datasets, Models, and a Research Roadmap

arXiv:2405.06818v2 Announce Type: replace Abstract: Natural Language Processing (NLP) for Ghana's 73 living indigenous languages remains deeply fragmented, under-resourced, and heavily skewed toward...

By Sheriff Issaka, Erick Rosas Gonzalez, Colene Agbo, Evans Kofi Agyei, Shruti Tyagi, John Emeka Eze, Enock Appiah Tieku, Junlin Fang, Thanh Do Nguyen, Juliet Arthur, Zhaoyi Zhang, Mihir Heda, Keyi Wang, Yinka Ajibola, Rebecca Akpanglo-Nartey, Frank Lawrence Nii Adoquaye Acquaye, Dennis Owusu, Jerry John Kponyo, Stephen Moore, Isaac Wiafe, Sean Du
arXiv Computer Vision
Sep 17

Think Before You Move: Latent Motion Reasoning for Text-to-Motion Generation

The paper introduces Latent Motion Reasoning (LMR), a two‑stage approach that separates text‑to‑motion generation into a planning phase and an execution phase. LMR uses a Dual‑Granularity Tokenizer to create a compressed, semantically rich reasoning latent for global trajectory planning and a high‑frequency execution latent for detailed motion fidelity. Experiments on T2M‑GPT and MotionStreamer show that this architecture improves both semantic alignment and physical plausibility compared to direct translation methods.

By Yijie Qian, Juncheng Wang, Yuxiang Feng, Chao Xu, Wang Lu, Yang Liu, Baigui Sun, Yiqiang Chen, Yong Liu, Shujun Wang
arXiv AI
Sep 17

A visual large language foundational model for medical image recognition using clinician-contributed online resources

The paper introduces ThoughtMed-1M, a large-scale medical visual question answering dataset built from de‑identified images and clinician‑generated commentaries, designed to capture structured clinical reasoning and image‑text alignment. Using this dataset, the authors train FOLTMed, a foundational large language model that achieves state‑of‑the‑art performance on 42 medical VQA benchmarks, with a macro accuracy of 85.4% and improved factuality and similarity metrics over existing models.

By Lingxuan Hou, Yuhua Xie, Yue Hu, Yan Zhuang, Junqi Li, Chengzhi Xia, Binh Phu Nguyen, Abubakar Siddique, Minh Nguyen, Yao Hou, Yanju Bao, Kexin Liu, Ke Chen, Jianjun Sun, Zeqi Li, Trung Nguyen, Jiangli Lin
arXiv Computer Vision
Sep 17

Understanding Dynamic Scenes at Gigapixel Scale: Wide-Area Spatio-Temporal Perception from UAVs

The paper introduces the Wide-area Spatio-temporal Scene Understanding (WSTU) problem, which demands simultaneous wide-area coverage, per-target resolution, and temporal continuity—capabilities lacking in existing datasets. To address this, the authors present HARD, an ultra‑high‑resolution (12768×9564) UAV dataset annotated for object detection, multi‑object tracking, and scene‑level visual question answering. They also propose a latency‑aware metric, streaming‑HOTA (s‑HOTA), and show through baseline experiments that high resolution and processing latency significantly impact detection, tracking, and VQA performance, revealing gaps in current methods for WSTU.

By Yuhang Zhu, Meiyi Zhu, Yunkai Dang, Zhangnan Li, Yuxuan Wang, Wenbin Li, Hongbing Pan
arXiv Computation and Language
Sep 16

TIAO: Token Importance-Aware Policy Optimization for Text Summarization

The paper introduces Token Importance-Aware Policy Optimization (TIAO), a reinforcement learning approach that improves text summarization by weighting token importance based on token dependency. TIAO reweights a trajectory’s advantage according to the overall dependencies of core tokens, addressing the limitation of previous methods that treat all tokens equally. Experiments demonstrate that a 7B foundation model enhanced with TIAO achieves performance comparable to GPT‑4 and GPT‑5‑nano on real‑world datasets.

By Qixiu Li, Chenlong Bao, Xiang Zhu, Xiaoyong Li, Ruixin Cao, Shukai Chen, Zhenxiong Zhou
arXiv Computation and Language
Sep 16

Spurious Tool Use: When RL Agents Learn the Wrong Reason to Act

The paper investigates how reinforcement learning can cause large language model agents to adopt shortcut policies for tool use, relying on superficial prompt cues rather than actual task needs. By creating synthetic environments that mix factual QA and math reasoning, the authors show that agents often invoke tools when cues are present, even when those tools are unnecessary, with spurious invocation rates rising up to 39%. They find that shortcut learning occurs mainly when agents have already mastered the target tool and that semantic alignment between cues and tools amplifies the effect. To counter this, they propose a dense, decision-level reward where an LLM judge assesses tool necessity, which reduces cue-driven tool use while maintaining performance.

By Yiwei Yang, Haoxiang Zhang, Bingbing Wen, Yao Lu, Yuchen Wu, Lei Zhang, Julian McAuley, Pan Lu, Bill Howe