Natural language processing

Classical and neural NLP: translation, question answering, tokenization and the evaluation of language understanding.

2,641 stories · RSS feed

arXiv Computation and Language
Aug 31

MemoryCard: Topic-Aware Multi-Modal Clue Compression for Long-Video Question Answering

MemoryCard is a video-memory-based augmentation framework designed to improve long-video question answering for Vision‑Language Models. It segments lengthy videos into semantically coherent units—each representing a distinct topic or event—by performing a self‑reading process over the video and aligned utterances. For each unit, the framework generates an event‑level video gist and selects representative visual moments, which are compiled into unified Memory Cards that are used for retrieval and answering questions, yielding up to a 21.8% relative accuracy improvement under comparable visual‑token budgets.

By Qing Yang, Pengcheng Huang, Xinze Li, Zhenghao Liu, Yukun Yan, Yu Gu, Ge Yu, Gang Li, Maosong Sun
arXiv Machine Learning
Aug 31

Not to Break, but to Attest: Adversarial Probes for Privacy-Preserving LLM Verification

The paper introduces a privacy‑preserving zk‑SNARK audit framework that uses adversarial‑style probes to detect logit drift between an approved large language model and a modified deployment. It offers three probe families—token‑based (black‑box), embedding‑based (gray‑box), and stress probes (partial white‑box)—allowing users to balance sensitivity, access, and cost. Experiments across LLM architectures and GPU platforms show token‑based probes achieve the highest mean sensitivity while remaining practical in a black‑box setting, with Groth16 proving times scaling modestly from 1.02 to 1.78 seconds and constant proof size.

By Cameron Wilding, Mina Shaker, Fatemeh Ganji
arXiv Computation and Language
Aug 31

SURE-Challenge: Evaluating Speech Evidence Before Speech-LLM Generation

The paper introduces the Speech-Unsupported Rejection Evaluation Challenge (SURE‑Challenge), a benchmark designed to test whether speech‑LLMs should accept or reject audio inputs before generating answers. Using LibriSpeech‑derived transcriptions paired with first‑word question answering, the authors evaluate various noise and silence conditions, and compare a simple energy‑plus‑Whisper‑score rule against a Qwen2‑Audio front‑end. On a 474‑row test set, the rule rejects 196 of 204 unsupported inputs while preserving accuracy on supported data, revealing a pre‑generation error mode that answer‑only scoring misses.

By Mengzhe Geng
arXiv Computation and Language
Aug 31

XHotpotQA: A Benchmark for Cross-Lingual Knowledge Composition in Multi-Hop Question Answering

XHotpotQA is a new benchmark for cross‑lingual knowledge composition in multi‑hop question answering. It presents each instance as an evidence‑dependency graph with explicit language assignments for the question, bridge evidence, answer‑bearing evidence, and distractors, and includes 15,661 training and 7,405 validation examples with sentence‑level support supervision. The dataset reveals significant performance drops when evidence spans language boundaries, providing a diagnostic tool for systems that must integrate evidence across languages.

By Iman Barati, Arash Ghafouri, Behrouz Minaei-Bidgoli
Towards Data Science
Aug 29

RAG Is Not the Whole Toolkit: The NLP Techniques Real Problems Still Need

The article argues that Retrieval-Augmented Generation (RAG) is only one tool in NLP, and many real-world problems—such as request classification, free‑text matching, table reading, and OCR noise cleaning—are better served by simpler, cheaper techniques. It emphasizes the importance of selecting the appropriate method for each task and highlights the engineering challenge of knowing which technique to apply.

By Kezhan Shi
arXiv Machine Learning
Aug 28

Stack Trace-Based Crash Deduplication with Transformer Adaptation

Stack Trace-Based Crash Deduplication with Transformer Adaptation introduces dedupT, a transformer‑based method that models entire stack traces instead of isolated frames. The approach first fine‑tunes a pretrained language model on stack traces and then trains a fully‑connected network to rank duplicate crashes. Experiments on four public datasets show dedupT improves Mean Reciprocal Rank by over 15% versus the best deep‑learning baseline and up to 10% over traditional methods, while also achieving higher ROC‑AUC for unique crash detection.

By Md Afif Al Mamun, Gias Uddin, Lan Xia, Longyu Zhang
arXiv Computation and Language
Aug 28

Cascaded Batch Prompting

Cascaded Batch Prompting introduces a two‑stage method that separates complex reasoning from symbol grounding to address the unpredictability of conventional batch prompting. Experiments on multiple‑choice question answering and natural language inference show that this approach outperforms standard single prompting while maintaining a speedup proportional to batch size. The technique establishes a new state‑of‑the‑art position on the Pareto frontier for efficiency and performance.

By Sho Hoshino, Peinan Zhang
arXiv AI
Aug 28

Beyond Factual QA: Mentorship-Oriented Question Answering over Long-Form Multilingual Content

The paper introduces MentorQA, a multilingual dataset and evaluation framework for mentorship-oriented question answering derived from long‑form videos. It contains nearly 9,000 QA pairs across four languages and defines evaluation dimensions such as clarity, alignment, and learning value that extend beyond factual accuracy. Experiments show that Multi‑Agent QA pipelines outperform other architectures, especially on complex topics and low‑resource languages, while automated LLM‑based evaluation shows variable alignment with human judgments.

By Parth Bhalerao, Diola Dsouza, Ruiwen Guan, Oana Ignat
arXiv AI
Aug 28

Co-Evolving Structured Knowledge and Reasoning in Language Models

The paper introduces KBevo, a co‑evolving framework that simultaneously builds a structured knowledge base and performs reasoning over it for knowledge‑intensive question answering. By optimizing both components end‑to‑end with QA outcome rewards, the system improves the quality and connectivity of the knowledge base, leading to higher answer reachability and better compositional factual reasoning. Compared to standard retrieval baselines, KBevo offers greater controllability and improved factual accuracy.

By Ryan Thomas Noonan, Linxi Zhao, Menghan Xu, Akanksha Sarkar, Mihir Mishra, Dongyoung Go, Kilian Q. Weinberger, Yoav Artzi, Jennifer J. Sun
arXiv AI
Aug 28

A Table Is Worth 64 Tokens: Pixel-level Compression for Multi-Table Document Question Answering

The paper investigates pixel-level table compression for document question answering, comparing five vision‑language models across two benchmarks and varying visual‑token budgets. It finds that representing tables as native‑resolution images matches text in performance and efficiency, while highly downscaled images still allow the model to identify relevant tables but lose readability, leading to longer reasoning traces. A two‑step, training‑free method first selects relevant tables from compressed images and then reasons over them at native resolution, saving 41% of tokens and improving accuracy by 7 points over single‑step native‑resolution QA, while using 15% fewer tokens than the most efficient single‑step compressed setup without accuracy loss.

By I\~nigo Alonso, Mirella Lapata
arXiv AI
Aug 28

Multivariate Diffusion Transformer with Decoupled Attention for High-Fidelity Mask-Text Collaborative Facial Generation

The paper introduces MDiTFace, a diffusion transformer designed for high‑fidelity mask‑text collaborative facial generation. It unifies tokenization of semantic masks and text, enabling synchronous multimodal feature interaction via stacked multivariate transformer blocks. A novel decoupled attention mechanism separates dynamic and static computations, allowing caching of static features and reducing mask‑condition overhead by over 94% while preserving performance, leading to superior facial fidelity and conditional consistency compared to existing methods.

By Yushe Cao, Dianxi Shi, Xing Fu, Xuechao Zou, Haikuo Peng, Xueqi Li, Chun Yu, Junliang Xing
arXiv Computation and Language
Aug 28

RuleWeaver: Benchmarking Rule-Centered Scenario Reasoning for Large Language Models

RuleWeaver is a benchmark construction framework designed to evaluate large language models’ ability to reason over complex, rule‑centered scenarios. It begins with corpus‑derived IF‑THEN meta rules, expands them into more intricate rules, and composes these into scenario‑based QA instances. The benchmark assesses not only final answer correctness but also process‑level metrics such as rubric‑based answer quality, rule recall, and rule precision, revealing that current LLMs achieve only about 50% of the maximum rubric score on these tasks.

By Bohan Yu, Shi-Yang Li, Pengfei Cao, Jun Zhao, Kang Liu
arXiv Computation and Language
Aug 28

Reasoning about In-Context Samples for Machine-Translation

The paper proposes a fragment‑based reasoning framework for large language model–based machine translation. It extracts parallel source‑target fragments from retrieved similar examples and uses these fragments as intermediate reasoning traces to generate the final translation. Experiments with the Qwen3 model across six languages and multiple domains show that this approach outperforms standard k‑shot or basic drafting methods.

By Maxime Bouthors, Josep Crego, Fran\c{c}ois Yvon
arXiv Computation and Language
Aug 28

Case2Flow: Bridging Patient Cases and Guideline Flowcharts through Multimodal Retrieval

Case2Flow is a new task that retrieves the most relevant guideline flowchart for a given patient case from a collection of medical guideline documents. The authors created FlowAtlas, a curated corpus of 202 flowcharts extracted from 2,080 guidelines, and a pipeline that generates 1,911 aligned case‑flowchart pairs. Their evaluation shows that existing multimodal retrieval methods often overrely on keywords and spurious token‑patch matches, and they propose CRISP, a training‑free scoring method that improves Recall@1 by up to 18.71 percentage points and gains preliminary feasibility evidence from a blinded physician assessment.

By Jiale Wei, Yufan Chen, Alexander Jaus, Zdravko Marinov, Julian Friedrich, Simon Rei{\ss}, Jens Kleesiek, Rainer Stiefelhagen
arXiv Machine Learning
Aug 28

Can a Model Catch Its Own Hallucinations for Free?: Label-Free Doubt Signals Hold Their Own Against a Labelled Dataset for Abstention

The paper investigates whether a language model’s own confidence can replace labeled data for teaching it to abstain from uncertain answers. By fine‑tuning models with LoRA to answer only when their frozen confidence is high and to say “I’m not sure” otherwise, the authors show that this label‑free approach matches label‑supervised abstention tuning on short‑form factual QA. The method works across six open‑weight models (1B‑8B) and is effective except for confidently wrong facts, which the confidence signal cannot flag.

By Ali Asaria, Tony Salomone, Deep Gandhi