Natural language processing

Classical and neural NLP: translation, question answering, tokenization and the evaluation of language understanding.

2,641 stories · RSS feed

arXiv AI
Aug 26

SonarLLM: A Native Sonar--Optical Multimodal Large Language Model for Underwater Perception

SonarLLM is a multimodal large language model that treats sonar as a native perceptual modality, combining a sonar‑specific encoder, physics‑aware feature enhancement, and reliability‑aware hierarchical fusion to align acoustic structure with optical semantics. The authors introduce SonarBench, a benchmark covering recognition, counting, visual question answering, and captioning across sonar‑only, optical‑only, and fusion settings, enabling controlled measurement of cross‑modal complementarity. SonarLLM achieves 72.0% macro accuracy on sonar‑only tasks and 68.7% under fusion, outperforming baselines by significant margins and demonstrating increasing fusion gains as optical visibility degrades.

By Cong Su, longxuan ma, Ling Dong, Guofeng Tang, Weijie Yin, Haohui Chen, Zhengtao Yu
arXiv AI
Aug 26

Gated Activation Steering for Reducing Sycophancy & Hallucination in Medical Question Answering

The paper introduces Gated Activation Steering, an inference-time intervention that jointly mitigates hallucination and sycophancy in medical question answering. By learning separate steering directions from contrastive clinical pairs and applying them to specific attention heads, the method uses behavior‑specific gates to intervene only when needed. Experiments on EHR‑based clinical questions show that the 4‑billion‑parameter model with gated steering outperforms its unsteered counterpart and rivals larger models in resisting user pressure.

By Himanshu Tripathi, Subash Neupane, Shaswata Mitra, Sudip Mittal, Noorbakhsh Amiri Golilarz, Shahram Rahimi
arXiv Machine Learning
Aug 26

Transformer Accelerator (TFA): A Macro-Op INT8 Hardware Chip for Transformer Inference and Machine Translation

The Transformer Accelerator (TFA) is a synthesizable, parameterizable INT8 memory‑to‑memory engine designed for transformer inference and machine translation. It features a one‑time‑multiplexed datapath that handles prompt processing and autoregressive generation, and implements key operations such as matrix multiplication, softmax, RMSNorm, and elementwise functions through eight 512‑bit macro‑op descriptors. In extensive verification, TFA achieved zero mismatches across 25 tests and 34 constrained‑random runs, matched floating‑point references on multiple translation tasks, and delivered a 20× speedup over a 22‑thread CPU while projecting significant energy reductions in larger designs.

By Shashank
arXiv Computer Vision
Aug 26

RT-NeuS: Towards Real-Time Neuro-Symbolic Video Understanding via Adaptive Temporal Verification

RT-NeuS is a neuro‑symbolic framework for long‑form video question answering that retains the accuracy and formal guarantees of temporal‑logic‑guided methods while dramatically reducing inference latency. It achieves this by using coarse‑to‑fine adaptive sampling to focus on query‑relevant frames and batched proposition detection with KV‑cache reuse, enabling all propositions to be evaluated in a single forward pass. Experiments on LongVideoBench, Video‑MME, and MLVU show up to a 13× speed‑up on an NVIDIA H200 GPU while matching or surpassing prior neuro‑symbolic accuracy.

By Shawn Liang, Sahil Shah, Chengwei Zhou, S P Sharan, Harsh Goel, Arnab Sanyal, Sandeep Chinchali, Gourav Datta
arXiv Computer Vision
Aug 26

AffineTok: Semantic Affine Consistency for Diffusion-Friendly Visual Tokenizer

arXiv:2608.23864v1 Announce Type: new Abstract: Visual tokenizers increasingly inject semantic supervision into latent spaces to make downstream diffusion easier. Yet how these semantics should be or...

By Junqiu Yu, Pandeng Li, Yikai Wang, Jiaxing Zhao, Yujie Wei, Kaixun Jiang, Quanhao Li, Hongtao Yu, Zhihang Liu, Zhaohe Liao, Junjie Zhou, Yun Zheng, Yu Liu, Yanwei Fu
arXiv AI
Aug 25

Length-Adaptive Decoding for Masked Diffusion Machine Translation

The paper introduces Entropy-Valley (EV), a training‑free method for selecting target length in masked diffusion machine translation. EV evaluates candidate canvases by mean predictive entropy from all‑mask forward passes, choosing the length the model is best prepared to fill. Compared to a baseline that uses training‑corpus length statistics, EV recovers a substantial portion of the COMET‑22 gain across En→Zh, Zh→En, and En→De, and expert evaluation confirms adequacy improvements, especially for Zh→En.

By Yan Zhan, Mengkai Hou, Wanting Zhang, Zhijun Gao
arXiv AI
Aug 25

Towards a Densing Law for User Representation Learning at Billion-Scale Capacity

The paper introduces a User Behavioral Densing Law that quantifies how the minimum sufficient tokenization capacity scales with data size in user representation learning. A pilot study on a billion‑scale Alipay dataset shows raw data scaling bottlenecks and the benefits of tokenization, while theoretical analysis and experiments reveal an approximately linear relationship between the logarithms of tokenization capacity and input data size. Using this law, the authors develop ALGN, an adaptive variable‑length tokenization method that outperforms existing baselines across diverse data sources and downstream tasks.

By Bin Dou, Junru Zhang, Zhaoyi Yuan, Wuliang Huang, Letian Gong, Baokun Wang, Huan Li, Yu Cheng, Weiqiang Wang
arXiv Computation and Language
Aug 25

Beyond Factual Knowledge: Benchmarking and Learning Step-Level Procedural Rule Reasoning in Large Language Models

arXiv:2608.22753v1 Announce Type: new Abstract: Large language models (LLMs) excel at text understanding and generation, yet still struggle to reliably understand and apply externally provided proced...

By Bohan Yu, Pengfei Cao, Chen Han, Chenxi Zhou, Zhiheng Zhang, Zhiyang Xie, Wenhao Teng, Xiangwen Liao, Jun Zhao, Kang Liu
arXiv Computation and Language
Aug 25

LuxIT: A Luxembourgish Instruction Tuning Dataset from Monolingual Seed Data

LuxIT is a monolingual instruction‑tuning dataset for Luxembourgish, created by synthesizing instruction‑answer pairs from native texts using the DeepSeek‑R1‑0528 model and a quality‑assurance LLM‑as‑judge process. The resulting 227,507 high‑quality pairs were used to fine‑tune 14 LLMs (≤15 B parameters), yielding an average accuracy increase of +5.37 percentage points on standardized Luxembourgish proficiency exams and improvements in macro‑averaged F1 on nine of the fourteen downstream NLP tasks. These findings demonstrate that synthetic monolingual data can effectively enhance LLM performance in low‑resource languages and reveal the complex relationship between exam performance and practical NLP gains.

By Julian Valline, Cedric Lothritz, Siwen Guo, Jordi Cabot