Natural language processing

Classical and neural NLP: translation, question answering, tokenization and the evaluation of language understanding.

2,601 stories · RSS feed

Hugging Face Trending Papers
Sep 3

Typological Feature Prediction with Large Language Models: An In-Context Learning Approach

The paper explores how large language models (LLMs) can predict typological features using an in-context learning approach with data from URIEL+ and Glottolog. Zero‑shot prompting alone is inadequate, but providing phylogenetic and geographic neighbour evidence enables LLMs to outperform all baselines, even for low‑resource languages. Additionally, most LLM rationales align with the supplied evidence, suggesting a move toward explainable typological predictions.

Hugging Face Trending Papers
Sep 3

KnowVis: Knowledge-Centric Visual Summarization for Video Lectures

KnowVis is a framework that converts linear video lectures into knowledge‑centric visual summaries. It first builds a detailed concept map from multimodal video content to identify key and challenging concepts, then organizes these into structured knowledge units before synthesizing engaging visual narratives. The authors also provide a dataset of 125 educational videos with 1,079 visual summaries and show through automated metrics and a human study that KnowVis outperforms existing methods in accuracy, clarity, and learning outcomes.

Hugging Face Trending Papers
Sep 3

The Impact of Synthetic Data Augmentation on Discourse-Pragmatic Function Classification

The study examines how the geometric placement of synthetic data affects discourse-pragmatic function classification. Using 410 annotated instances of the word "look" and synthetic examples generated by Llama 3.1, the authors partitioned the synthetic data by cosine distance in RoBERTa space and compared six training conditions. While all augmentation conditions improved macro‑F and accuracy over a real‑only baseline, the nearest synthetic examples (NEAR) yielded the largest macro‑F gain (0.113) and a distance‑balanced mix achieved the highest accuracy (0.748), though none improved AUC.

Hugging Face Trending Papers
Sep 3

When Retrieval Helps: Selective Retrieval for Single-Turn Mental-Health QA

The paper investigates how retrieval-augmented generation (RAG) affects single‑turn mental‑health question answering. It finds that always using retrieval improves response specificity but can lower overall quality and increase safety‑related failures. A lightweight selective retrieval policy, guided by psychoeducational, coping, and safety needs, better balances these trade‑offs by activating retrieval only when necessary.

arXiv AI
Sep 3

A Survey of Transformer-based Language Models with Focus on Efficiency

The paper surveys Transformer-based large language models (LLMs) with a focus on efficiency, reviewing 312 articles that cover data curation, model design, downsizing, and dynamic inference. It also examines efficiency in adaptation strategies such as pre‑training, fine‑tuning, prompt‑engineering, and Retrieval‑Augmented Generation (RAG). A statistical analysis and evaluation of over 30 prominent NLP models on 13 benchmarks provide insights into both efficiency and efficacy, highlighting trends toward sustainable NLP practices.

By Wazib Ansar, Saptarsi Goswami, Amlan Chakrabarti
arXiv AI
Sep 3

Direct Construction of Disambiguated Knowledge Bases from Large Language Models

The paper introduces GPTKB 2.0, a method for building disambiguated knowledge bases directly from large language models. It addresses the lack of native entity representation in LLMs by performing on‑the‑fly disambiguation of entities, relations, and classes, achieving a million‑scale KB with over 1 million disambiguated entities and 38.4 million triples. The authors analyze trade‑offs among accuracy, scale, and cost, and release the system at https://gptkb.org/.

By Yujia Hu, Tuan-Phong Nguyen, Simon Razniewski
arXiv Computation and Language
Sep 3

MUDIDI: A Two-Stage Framework for Multilingual Dictionary Digitization with Language Models

MUDIDI is a two-stage framework designed to digitize multilingual dictionaries that are currently only available as scanned images. The first stage assesses character recognition and markup preservation, while the second stage segments dictionary entries and maps them into the SIL Multi-Dictionary Formatter schema. The authors also release a dataset of 30 annotated dictionaries and benchmark OCR, LLM, and VLM systems, finding that LLMs generally outperform others and that providing additional context improves digitization quality.

By David Setiawan, Temuulen Khishigsuren, Milind Agarwal, Pagnarith Pit, Aso Mahmudi, Ekaterina Vylomova
arXiv Machine Learning
Sep 3

Oracle, will I ever learn? A study of prediction convergence and complementarity across link prediction models

The paper investigates how different link prediction models for knowledge graphs produce varying predictions and explores the extent of their complementary knowledge. By evaluating an oracle that selects the best prediction from a set of models, the authors show a significant performance gap between individual models and the oracle, indicating substantial complementarity. However, this complementarity quickly saturates as more models are added, leaving many queries unsolved even with many models.

By Guillaume M\'erou\'e, Fabien Gandon, Pierre Monnin
arXiv AI
Sep 3

Response-free item difficulty modelling for multiple-choice items with fine-tuned transformers: Component-wise representation and multi-task learning

The paper proposes a response‑free method for estimating difficulty of reading‑comprehension multiple‑choice items by fine‑tuning a transformer on item wording. It introduces two extensions to a baseline joint‑encoding model: a component‑wise variant that encodes passage, question, and options separately, and a multi‑task variant that adds a question‑answering auxiliary task. Experiments on a corpus of nearly 30,000 items show that both extensions outperform the baseline, especially the multi‑task variant across all metrics and the component‑wise variant in rank ordering, even with limited training data.

By Jan Net\'ik, Patr\'icia Martinkov\'a
arXiv Computer Vision
Sep 3

GEM: Generating LiDAR World Model via Deformable Mamba

GEM is a Generative LiDAR world model that uses a deformable Mamba architecture to better handle the disorder of LiDAR point clouds and distinguish dynamic objects from static structures. The model tokenizes LiDAR sweeps, unsupervisedly disentangles dynamic and static features, and applies a tri‑path deformable Mamba for selective scanning and adaptive gating fusion, improving spatial‑temporal understanding. Experiments show GEM outperforms existing methods across multiple benchmarks, and it can be paired with a planner and BEV controller for autonomous rollout and "what‑if" scenario generation.

By Yang Wu, Zhaojiang Liu, Qiang Meng, Youquan Liu, Renliang Weng, Jianjun Qian, Jian Yang, Jin Xie
arXiv Computation and Language
Sep 3

Improving Health Literacy through Lay Summarization of Radiological Reports: An Evaluation of BioNER and Retrieval-Augmented Generation

The paper examines how Retrieval-Augmented Generation (RAG) and Named Entity Recognition (NER) affect the quality of lay summaries of radiology reports. Using a framework that extracts clinically relevant findings via NER and grounds them with RAG, the authors evaluate few‑shot and fine‑tuned versions of Qwen and BioBART. Results show that NER consistently improves readability and overall quality, RAG alone offers no benefit and can introduce hallucinations, and the best performance comes from fine‑tuned BioBART with NER.

By Egecan \c{C}elik Evgin, \.Ilknur Karadeniz, Olcay Taner Y{\i}ld{\i}z
arXiv Computation and Language
Sep 3

Grounded, Compute-Efficient LLM Policy Agents for Energy-Poverty Equity in Physically-Constrained Peer-to-Peer Energy Markets

The paper introduces EqGrid, a closed‑loop simulation that uses a low‑frequency, open‑weight LLM policy agent to set price, carbon limits, and subsidies for a community of empirically‑grounded household personas, while high‑frequency multi‑agent RL traders clear a continuous double auction on a physically constrained IEEE‑33‑bus grid. It demonstrates that the LLM can reduce energy‑poverty inequality—lowering the Gini of energy burden from 0.351 to 0.305 and mean burden by 28%—without increasing net grid cost, and that a compressed sub‑1B model retains 92–95% of this benefit at dramatically lower inference energy. The study also establishes a compute‑efficiency frontier and a decoupled‑safety design that eliminates grid violations. whyItMatters":"By showing that a lightweight LLM can effectively manage energy markets to reduce poverty and inequality while staying energy‑efficient, the work offers a practical, low‑carbon AI solution for humanitarian energy‑poverty interventions."

By Kunal Jadhav, Siddhesh More
arXiv Computation and Language
Sep 3

CARPAS: Towards Content-Aware Refinement of Provided Aspects for Summarization in Large Language Models

The paper introduces CARPAS, a new task that dynamically refines user-provided aspects for aspect-based summarization in large language models (LLMs). It presents three new datasets and evaluates four prompting strategies, finding that LLMs tend to over-generate aspects, leading to overly long and misaligned summaries. To address this, the authors propose a two-stage framework that first generates lightweight scope guidance before aspect refinement and summarization, which improves focus, reduces over-generation, and enhances performance across all datasets.

By Yong-En Tian, Yu-Chien Tang, An-Zi Yen, Wen-Chih Peng
arXiv Computation and Language
Sep 3

Are Non-English Papers Reviewed Fairly? Language-of-Study Bias in NLP Peer Reviews

The paper investigates language-of-study (LoS) bias in NLP peer reviews, defining and distinguishing negative and positive forms of bias. Using a new dataset, LOBSTER, and an LLM-based detection pipeline, the authors analyze 15,645 reviews and find that non‑English papers experience significantly higher bias rates, with negative bias outweighing positive bias. They further identify four subcategories of negative bias, noting that demanding unjustified cross‑lingual generalization is the most common.

By Ehsan Barkhordar, Abdulfattah Safa, Verena Blaschke, Erika Lombart, Marie-Catherine de Marneffe, G\"ozde G\"ul \c{S}ahin
arXiv Computation and Language
Sep 3

Cite or Decline: A Strict Course-Grounded Chatbot for STEM Lecture Videos

The paper presents a semester-long deployment of the VideoPoints platform, featuring a retrieval‑augmented chatbot that answers questions using only the active course’s lecture materials and provides clickable, timestamped citations. Across 833 interactions, 70.5% of messages included citations, no cross‑course references were made, and the bot declined to answer when no lecture evidence matched. The study also shows that the system improves correct‑lecture retrieval by 6.3 percentage points over dense‑only retrieval on the EduVidQA benchmark.

By S M Masrur Ahmed, Jaspal Subhlok
arXiv Machine Learning
Sep 3

A Unified Rate-Distortion Perspective on Vector, Product, and Scalar Quantization

The paper introduces a unified rate–distortion framework for discrete visual tokenization, encompassing vector, product, and scalar quantization. It shows that minimizing distortion, rather than maximizing codebook utilization, is the key objective for reconstruction fidelity and establishes fairness conditions for comparing quantizers. Under these conditions, the study confirms the distortion hierarchy VQ–PQ–SQ and demonstrates that modern VQ methods achieve the lowest distortion.

By Xianghong Fang, Wenlong Mou, Yuan Yuan, Dehan Kong, Tim G. J. Rudner
arXiv Computer Vision
Sep 3

Spatially Aware World Action Model via Geometric Latent Diffusion

The paper introduces Spatially Aware World Action Model (SA‑WAM), a diffusion‑based framework that extends existing World Action Models by incorporating depth information alongside RGB to enable 3‑D‑aware action and future‑state prediction. SA‑WAM repurposes a pretrained video diffusion model, using a nonlinear encoding to map unbounded depth into the tokenizer’s bounded domain, thus preserving pretrained visual priors without 3‑D‑specific fine‑tuning. The model achieves state‑of‑the‑art performance on RoboCasa and LIBERO‑Plus benchmarks and demonstrates superior real‑world performance on a UR5 robotic arm in randomized environments, while also providing analysis linking world‑model prediction quality to rollout success.

By Javier Alejandro Lopetegui Gonzalez, Paul Pacaud, Cordelia Schmid
arXiv Computation and Language
Sep 3

The Geometry of LLM-as-Judge: Why Inter-LLM Consensus Is Not Human Alignment

The paper investigates whether consensus among large language model (LLM) judges truly reflects human alignment. By treating each judge’s scores as vectors, the authors measure spread, effective rank, and angles to human scores across 42 judges on Indic benchmarks, revealing that inter‑judge agreement often mirrors shared blind spots rather than human judgments. They find that while judges agree as much as humans, they only reach 58‑66% of human agreement and frequently focus on axes humans do not weight, indicating that ensemble agreement alone is insufficient evidence of alignment.

By Sourabrata Mukherjee, Hamna Hamna, Kalika Bali, Sunayana Sitaram
arXiv AI
Sep 3

InfraPatch: Cross-Task Targeted Grayscale Patch Attacks on Infrared-Adapted Vision-Language Models

InfraPatch is a white‑box, per‑instance framework that generates small grayscale patches to target infrared‑adapted vision‑language models (IR‑VLMs). The method optimizes a single‑channel patch within a 5% local‑area budget, using proxy‑guided placement and task‑adaptive objectives to induce desired behaviors in image classification, captioning, and binary visual question answering. Across ten IR‑VLM variants tested on synthetic infrared images, InfraPatch achieves targeted attack success rates ranging from 86% to 100%, revealing significant vulnerability differences among architectures and tasks.

By Chengyin Hu, Dingyi Lu, Jiaju Han, Xiang Chen, Weiwen Shi, Jiahuan Long, Yiwei Wei, Jiujiang Guo
arXiv AI
Sep 3

APEx: Distillation of Agent Procedural Experience for Adaptive Deep Research Question Answering

APEx is a hierarchical framework that organizes a deep research agent’s interaction history into instance-level trajectory memories and category-level procedural skills. It couples these through an Executor, Distiller, and Planner, trained with a three-stage alternating GRPO paradigm to enable reward-guided skill distillation. At test time, distilled skills act as procedural priors for online Planner adaptation via skill-guided reinforcement learning, achieving state‑of‑the‑art results on seven benchmarks, outperforming GPT‑5.4 by 14.7 points and the best memory‑augmented baseline by 3.0 points.

By Jie Ding, Rui Sun, Xinyuan Zhang, Zeyu Zhang, Xin Liu