Natural language processing

Classical and neural NLP: translation, question answering, tokenization and the evaluation of language understanding.

2,564 stories · RSS feed

arXiv Machine Learning
Sep 18

The Environmental Impacts of Language Model Training Keep Rising Now is the Time to Catch Impacts on the Rebound

The paper analyzes the environmental footprint of machine learning model training, focusing on large language models and their hardware. It finds that energy use and environmental impacts have risen exponentially over the past decade, even when employing carbon‑efficient electricity and more efficient hardware. The study argues that optimization strategies alone cannot curb these impacts due to a rebound effect, and stresses the need to evaluate hardware life‑cycle impacts and integrate environmental metrics into NLP research practices.

By Cl\'ement Morand (STL), Anne-Laure Ligozat (ENSIIE, LISN, STL), Aur\'elie N\'ev\'eol (STL, LISN)
arXiv Computation and Language
Sep 18

Less Is More: Graph-free Multimodal RAG via Multi-signal Late Fusion

The paper introduces TrioRAG, a graph-free multimodal retrieval-augmented generation framework that combines evidence from the question, an anchor image, and a VLM-enhanced query via late fusion. It also presents AutoQA, a benchmark featuring noisy web-sourced images that require reasoning across manuals. TrioRAG outperforms graph-based systems on three benchmarks while cutting costs and speeding up inference by 1.6–2.3×.

By Tithi Rakshit, Hongkuan Zhou, Lavdim Halilaj, Yuqicheng Zhu
arXiv AI
Sep 18

M2Tok: Multi-head Multi-codebook Discrete Action Tokenization for Vision-Language-Action Models

The paper introduces M2Tok, a Multi-head Multi-codebook Action Tokenizer that reduces reconstruction loss for continuous action signals by decomposing latent features into multiple heads and assigning independent codebooks to each. This design expands representational expressivity, leading to lower reconstruction error and higher success rates in Vision‑Language‑Action models evaluated on RoboTwin, Simpler‑Env, and zero‑shot real‑world tasks. Ablation studies confirm the effectiveness of both multi‑head and multi‑codebook mechanisms.

By Chunpu Xu, Zhixuan Liang, Yuhao Zhang, Chi-Min Chan, Jessie Wang, Yang Xiao, Mengkang Hu, Xiaokang Yang, Yao Mu
arXiv AI
Sep 18

How to Guide Your Language Flow

The paper introduces probe guidance, a technique that leverages frozen internal states of a diffusion model to generate a guidance signal without requiring an extra forward pass during inference. This method improves continuous diffusion language models, achieving state‑of‑the‑art results on unconditional generation and enhancing performance on multiple‑choice question answering for a 1.7B model. The authors also use probes to analyze autoguidance, revealing that the weak model must originate from a low‑entropy training region to align dynamics with the strong model.

By Rohit Dilip, Tianrong Chen, Yuyang Wang, David Van Valen, Joshua Susskind, Miguel Angel Bautista
arXiv Computation and Language
Sep 18

Schema-Anchored Latent Reasoning for Semantic Parsing-Based Knowledge Base Question Answering

The paper introduces SALR, a schema‑anchored latent reasoning approach for generating logical forms in knowledge‑base question answering. SALR delays explicit schema commitments by generating continuous thoughts in hidden states and aligns these thoughts with a codebook of KB schema elements, guided by an alignment objective derived from gold logical forms. Experiments on GrailQA and WebQSP demonstrate that SALR consistently outperforms strong baselines, notably improving compositional question performance by 2.86 F1 points over TIARA.

By Guangze Gao, Zixuan Li, Sikui Zhang, Chunfeng Yuan, Wenjuan Li, Bing Li, Xiaolong Jin, Weiming Hu
arXiv AI
Sep 18

AntiGrounding: Executable Robot Trajectories as Visual Prompts for VLM-Guided Manipulation

AntiGrounding is a visual action-selection framework that turns short robot trajectories into both executable motion plans and rendered prompts for vision‑language model evaluation. After filtering for feasibility, each trajectory is scored on safety, task alignment, efficiency, and physical plausibility using structured multi‑view visual question answering, and the best trajectories are refined and validated by a digital twin before real‑world execution. In eight real‑world manipulation tasks, the system achieved a 71.25% success rate with a single GPT‑6 Astra evaluator, outperforming baseline methods.

By Wenbo Li, Yiteng Chen, Wenhao Li, Qingyao Wu
arXiv Machine Learning
Sep 18

VisKG-LM: Compiling Knowledge Graphs into Visual Memory for Multiple-Choice Question Answering

VisKG‑LM proposes compiling retrieved knowledge graph subgraphs into static visual memories rather than re‑encoding them during each inference step. The method serializes each subgraph as Relation‑Labeled Paths, renders them as images that preserve the graph’s branching structure, and caches these images for reuse. At inference, a language model processes the question and candidate text first, then consults the cached visual memory only at its final layer, yielding improved performance on CommonsenseQA, OpenBookQA, and MedQA‑USMLE compared to both text‑only baselines and a large vision‑language model.

By Yixin Peng, Er Jin, Shiwei Luo, Diego Collarana, Stefan Decker
arXiv Machine Learning
Sep 18

The Life of a Token: from Words to Bits on the Wire

The article "The Life of a Token: from Words to Bits on the Wire" explores how large language models convert text into network traffic during training. It traces the transformation from words to tokens, then to vectors, and finally to binary streams that traverse high‑performance computing systems. Using Dante’s Divine Comedy as a case study, the tutorial examines how tokenization, embeddings, and parallelization affect the volume, structure, and timing of data exchanged across the network, and provides analytical traffic models and numerical examples to clarify the communication demands of LLM training.

By Davide Avesani (CEDRIC - ROC), Pengwenlong Gu (CEDRIC - ROC), Sotiris Skaperas (Cnam), Stefano Secci (CEDRIC - ROC)
arXiv AI
Sep 18

Chunk Twice, Embed Once: A Systematic Study of Segmentation and Representation Trade-offs in Chemistry-Aware Retrieval-Augmented Generation

The study investigates how document segmentation and chunk representation affect retrieval-augmented generation (RAG) for chemistry texts. Using the ChemQuests corpus, the authors benchmark 41 embedding models and evaluate them across five chunking strategies, seven chunk sizes, and various overlap settings. They find that embedding choice has the largest impact, with models like E5, BGE, and Nomic performing best, and recommend medium-to-large chunks with fixed-token, recursive-token, or hierarchical-section chunking and low overlap for effective chemistry-aware RAG.

By Mahmoud Amiri, Thomas Bocklitz
arXiv AI
Sep 18

Not All AI Agents Are Equal: Characterizing Resource and Performance Dynamics

The paper investigates how large‑language‑model (LLM) based AI agents mix latency, local resource usage, and container bottlenecks when processing user requests that involve remote LLM calls and local tool execution. By measuring three representative tasks—retrieval‑augmented question answering, web search, and software coding—the authors show that agents exhibit diverse resource dynamics, with concurrent requests revealing task‑specific bottlenecks in CPU, disk I/O, and memory. Leveraging these insights, they propose CPU‑aware tool admission and task‑aware CPU allocation, achieving up to a 5.4× speed‑up for CPU‑sensitive tasks and a 32% reduction in average latency across multiple tasks.

By Wonmi Choi, Minuk Park, Zhixiong Niu, Yongqiang Xiong, Chuck Yoo, Gyeongsik Yang
arXiv AI
Sep 18

An Empirical Study of Harness Design for Coding Agents

The study investigates how individual components of a coding harness—planning, action space, and context management—affect autonomous coding agents’ performance. By fixing the execution loop and varying these components across 176 settings on SWE‑Bench Verified and Terminal‑Bench 2.1, the authors find that context management is most valuable when context windows are tight, staging rule‑based elision before LLM summarization yields the best efficiency, planning serves as an accuracy scaffold for weaker models and a cost saver for stronger ones, and predefined tools help models with limited bash skills while bash‑capable models benefit from a bash‑only interface. Trajectory‑level analysis shows that context management lengthens execution paths, planning alters where trajectories terminate, and the action space determines code granularity, offering a modular framework for future harness design.

By Run-Ze Fan, Zihao Zhang, Simin Ma, Yebowen Hu, Shouju Wang, Kaiqiang Song, Fei Liu, Hamed Zamani, Xiaoyang Wang
Hugging Face Trending Papers
Sep 17

DocAttriBench: Benchmarking Answer Grounding in Document Visual Question Answering

DocAttriBench (DAB) is a large-scale benchmark that provides fine-grained, element-level source attribution for Document Visual Question Answering (VQA). It introduces MAPPET, a Mask-based Perplexity-Derived Attribution method that uses document layout and language modeling to identify the most informative layout element for each answer. The benchmark contains 237k documents and 296k question-answer pairs with grounding annotations, and it evaluates multimodal LLMs on answer accuracy, attribution accuracy, and overall answer quality.

Hugging Face Trending Papers
Sep 17

VākQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question Answering

VākQA is a newly introduced benchmark for Telugu spoken factoid question answering, comprising 2,001 question‑answer pairs across six domains, 2.53 hours of speech audio, bilingual transcriptions, and human‑verified reference answers. The study validates automatic evaluation methods against human judgments, finding that Gemini‑as‑a‑judge best approximates human ratings but is inconsistently strict, while open‑weight judges tend to penalize correct Telugu answers that differ in surface form. Using this validated setup, the authors benchmark proprietary and open‑weight models, highlighting challenges such as cultural specificity loss in translation, phonetic confusions from speech input, and compounded errors from cascaded ASR‑MT pipelines.

arXiv Computation and Language
Sep 17

Divide and Conquer: Mixture-of-Bottleneck Experts in Informative Ordinal Space for Video-based Multimodal Sentiment Analysis

The paper proposes a Mixture-of-Bottleneck (MoB) framework for video-based multimodal sentiment analysis that treats sentiment as an ordinal regression problem, splitting it into polarity recognition and intensity prediction. MoB assigns modality‑specific latent experts to each sub‑task, learns compact, task‑relevant representations via an information bottleneck, and fuses these experts with a multimodal bottleneck routing module and hard mining strategy. Experiments on four datasets and language models demonstrate that MoB captures fine‑grained intra‑ and inter‑modal dynamics, improving performance and enabling more trustworthy localization of nuanced sentiment signals.

By Ronghao Lin, Qiaolin He, Zefeng Lu, Yichu Liu, Li Huang, Sijie Mai, Haifeng Hu, Yap-peng Tan
arXiv Computer Vision
Sep 17

Sim-to-Real Traffic Scene Understanding by Decoupling Semantics from Caption Generation with V-JEPA

The paper presents a decoupled framework for sim-to-real traffic scene understanding, separating semantic fact extraction from caption generation. It uses a frozen V-JEPA encoder for predictive scene representations and a lightweight Llama-based predictor for VQA, followed by a training-free structured refinement that leverages statistical priors, inter-question relationships, and temporal consistency. The refined facts are then fed to Qwen3-VL-8B to produce pedestrian and vehicle descriptions, achieving top performance on the 2026 AI City Challenge Track 2 benchmark with 87.09% VQA accuracy and an overall S2 score of 60.0853.

By Nguyen Hoai Thuong Bui, Thanh Nguyen Vo, Trinh Tra Giang Nguyen, Ha Duc Bui