Natural language processing

Classical and neural NLP: translation, question answering, tokenization and the evaluation of language understanding.

2,564 stories · RSS feed

arXiv Machine Learning
Sep 23

Geometry-Aware Hyperbolic Residual Quantization

The paper introduces Geometry‑Aware Hyperbolic Residual Quantization, a method that adapts residual vector quantization to hyperbolic space while preserving its telescoping structure. It achieves this by using Hyperbolic Residual Aggregation in the forward pass and a discounted Hyperbolic Straight‑Through Estimator in the backward pass, thereby avoiding geometric inconsistencies and unstable gradients. Experiments on hierarchical prediction, recommendation, image tokenization, and neural audio coding demonstrate improved stability and structural organization of hyperbolic residual codes, with a noted trade‑off between compression efficiency and hierarchical organization.

By Alessio Colombo, Melika Ayoughi
arXiv Machine Learning
Sep 23

What Does Chain-of-Thought Entropy Measure? A Channel Audit of Scaffolding, Routing, and Content

The paper investigates how entropy over chain‑of‑thought tokens influences policy decisions such as gradient application, pruning, and collapse detection. By separating scaffold tokens from substantive content, the authors analyze entropy, Kullback–Leibler divergence, and entropy velocity for each channel, proving differences between raw and content conventions and bounding answer diversity. Empirical results across 23 configurations show that scaffold tokens can account for up to 41% of high‑entropy positions, with entropy share growing through distillation, while content conventions outperform raw surprisal on compression tasks and reveal significant answer leakage in re‑fed chains.

By Marios Papamichalis, Regina Ruane
arXiv Machine Learning
Sep 23

Beyond Reconstruction Error: Analytical and Data-Driven Action Tokenization for Autoregressive Vision-Language-Action Models

The paper investigates how different action tokenization methods affect closed‑loop control in autoregressive vision‑language‑action models. It compares analytical, linear, and nonlinear representations, showing that lower reconstruction error does not guarantee better policy performance. The study highlights the need to evaluate tokenization on multiple criteria, including sequence predictability and decoder stability, rather than relying solely on reconstruction fidelity.

By Yuxin Yang, Gaohan He, Changxue Guan, Hangming Liu
arXiv Computation and Language
Sep 23

Retrieved-Span Training for Efficient Query-Focused Meeting Summarization on QMSum

The paper introduces a retrieved‑span training approach for query‑focused meeting summarization on the QMSum benchmark. By fine‑tuning a 406 M Fusion‑in‑Decoder model on 2,000‑word retrieved spans, the authors recover a 6.30 ROUGE‑1 loss incurred when moving from capped long input and achieve a test score of 36.33 ROUGE‑1, comparable to a larger 1.2 B system. The smaller model uses roughly one‑third the parameters and less than half the peak inference memory, while span‑regime fine‑tuning adds significant gains over the baseline. "whyItMatters":"The study demonstrates that efficient, smaller models can match or exceed larger systems on QMSum using span‑based fine‑tuning, offering a practical path for scalable query‑focused meeting summarization."

By Edward Xi Yang (Ertas AI)
arXiv Computation and Language
Sep 23

ARAFA: An LLM-Generated Arabic Fact-Checking Dataset

A new large-scale Arabic fact‑checking dataset called Arafa has been created using an automated pipeline that generates claims from Arabic Wikipedia, mutates them into counterfactuals, and validates them against supporting or refuting evidence. The dataset contains 181,976 claim‑evidence pairs labeled as supported, refuted, or not enough information, and human evaluation shows high inter‑annotator agreement and strong validation accuracy. Fine‑tuned transformer models on Arafa achieve a Macro F1‑score of 77%, demonstrating its usefulness for Arabic fact‑checking tasks.

By Christophe Khalil, Shady Elbassuoni, Rida Assaf
arXiv Computation and Language
Sep 23

Calibration as a First-Class Criterion in LLM Evaluation

The paper argues that calibration—how well a language model’s confidence aligns with its actual correctness—should be a standard evaluation metric for large language models (LLMs). It notes that while calibration metrics exist, they are rarely applied outside specialized NLP subfields, leading to unverified confidence scores in new models, datasets, and benchmarks. The authors highlight the risks of miscalibration both at deployment (overconfident errors causing harm) and in research workflows (affecting LLM-as-a-judge, synthetic data generation, and active learning). They call for every NLP subfield to pair its primary performance metric with a calibration score, treating calibration as an essential property of every model.

By Mario Sanz-Guerrero, Katharina von der Wense
arXiv Computation and Language
Sep 23

Knowledge Pull Requests for Continual Document Authoring

The paper introduces Knowledge Pull Requests (KPRs), a framework that enables continual document authoring by making each change interpretable. KPRs extract claims from new knowledge sources, filter and route them to appropriate sections, and flag conflicts with existing content, producing a ChangeLog that separates knowledge changes from textual edits. Experiments on revising Wikipedia and updating query-driven reports show that KPRs integrate more information, better preserve existing content, and improve question answering performance compared to rewriting from scratch or using frontier models with search.

By Alexander Martin, Benjamin Van Durme
arXiv Computer Vision
Sep 23

Reading Right, Answering Wrong: How Visual Configuration Changes Affect Evidence Use in VLMs

Vision‑language models (VLMs) can lose accuracy when images are resized, even with minimal changes. The study shows that such small visual configuration changes—like tiling or token arrangement—cause more correctness flips across multiple checkpoints and benchmarks. Interestingly, in many cases the models still read the correct answer but fail to use it, and attention interventions reveal that configuration shifts weaken the use of readable information. By guiding models with field cues and their own transcriptions, the authors correct 97.2% of these errors.

By Dingyang Lin, Yingfeng Luo, Chenglong Wang, Chenwei Zhu, Anxiang Ma, Jingbo Zhu, Tong Xiao
arXiv Computer Vision
Sep 23

Identity-Centric Video Summarization via Hierarchical Fusion of Biometric, Appearance, and 3D Body Features

The paper introduces a video summarization method that fuses facial embeddings, 3D body‑shape features, and visual appearance within a multi‑object tracking framework. By hierarchically assigning identities and using bidirectional anchoring, it robustly recovers trajectories even under heavy occlusion or low visual quality. Keyframes are selected through a multi‑factor weighting scheme that balances biometric clarity, social interaction, and motion dynamics, while Adaptive Non‑Maximum Suppression guarantees temporal diversity, resulting in a compact, identity‑centric summary.

By Milad Mirjalili, Enrique Alegre Guti\'errez, Eduardo Fidalgo Fern\'andez, V\'ictor Gonz\'alez Castro, Roc\'io Alaiz Rodr\'iguez, Manuel Castej\'on Limas
arXiv Computer Vision
Sep 23

LLaVA-Assessor: Building the Foundation LMM For Visual Quality Assessment

LLaVA‑Assessor is a unified large multi‑modal model (LMM) designed for visual quality assessment, combining image and video inputs. It introduces a two‑task framework—quality interpretation and quality scoring—supported by an adaptive architecture, a rigorous human‑annotated dataset, and a machine‑synthesized data expansion pipeline. The model employs a prompt‑disentanglement strategy to stabilize multi‑task training and achieves strong performance across 11 quality scoring test sets and 4 interpretation benchmarks.

By Ziheng Jia, Zicheng Zhang, Jiaying Qian, Guangtao Zhai, Xiongkuo Min
arXiv Computation and Language
Sep 23

MultiViewDx: Evidence-Linked Multi-View Clinical Diagnosis

MultiViewDx is a physician‑validated multimodal instruction dataset that links medical imaging studies with patient context and normalizes heterogeneous reports into an evidence‑linked workflow (evidence → findings → differential discussion → diagnosis). The dataset covers a wide range of imaging modalities and uses a unified image‑text retriever to ensure that instruction synthesis is grounded in source‑supported evidence. Fine‑tuned models on MultiViewDx achieve the highest average accuracy on four MedVQA benchmarks and receive the strongest overall rating on JAMA Clinical Challenge cases, with ablations confirming the importance of case‑level multi‑view organization and evidence‑linked reasoning.

By Junda Wang, Zonghai Yao, Yujan Ting, Eric Z. Chen, Hieu Tran, Hong Yu, Weijing Huang, Terrence Chen
arXiv AI
Sep 23

ScholarStack: Layered Research Asset Orchestration and Cross-Task Reuse for Scientific Agents

arXiv:2609.23735v2 Announce Type: new Abstract: Scientific agents support a range of literature-based research tasks, such as retrieval, question answering, evidence-grounded generation, and claim as...

By ScholarSeed AI Team, Caoqinwei Gong, Xue Jiang, Wei Luo, Xiaoyu Qiu, Jiayi Sheng, Yi Wang, Zheng Yu, Ao Zhang, Haifan Zhang, Hanwei Zhang, Jihai Zhang, Yuan Cao, Wei Chen, Liyun Dai, Wenkai Fang, Guanglei Wang, Kai Ying, Tingyu Zhu, Wotao Yin