Natural language processing

Classical and neural NLP: translation, question answering, tokenization and the evaluation of language understanding.

2,564 stories · RSS feed

arXiv Machine Learning
1d ago

Universal Byte-Level Encoding: UTF-8/UTF-16 Routing to Reduce Cross-Script Token-Budget Disparities

The paper introduces Universal Byte-Level Encoding (UBE), a dual‑alphabet tokenizer that routes 3‑4‑byte UTF‑8 characters through UTF‑16 while keeping 1‑2‑byte characters on the UTF‑8 path. This design lowers the worst‑case token‑budget disparity for high‑premium scripts without increasing costs for efficient English spans, and it preserves standard BPE merges and exact decoding. In extensive Unicode audits and multilingual language‑model experiments, UBE matches or improves token‑count efficiency and context usability compared to traditional byte‑pair encoding.

By Hyunsik Kim, Youngmoon Jung
arXiv Machine Learning
1d ago

Why Does Train-Validation Separation Emerge? Update-Pressure Density Dynamics in Pretrained Backbones

The paper investigates why the train‑validation performance gap widens during fine‑tuning of pretrained models. It proposes a dynamic structural explanation: as training proceeds, updates shift from broadly reusable features to more example‑specific ones, increasing gradient heterogeneity and the gap. Experiments on synthetic ResMLP hierarchies, NLP models (RoBERTa, DeBERTa, Qwen) across six datasets, and vision models (ResNet‑18) confirm that higher reliance on private features correlates with larger accuracy gaps, supporting the proposed account.

By Yuchen Li, Mingyu Du, Zongqi Fan, Ken-Tye Yong, Nguyen H. Tran
arXiv Machine Learning
1d ago

Deep Symmetric Autoencoders from the Eckart-Young-Schmidt Perspective

The paper presents a theoretical analysis of symmetric autoencoders, a class of deep learning architectures frequently used in machine learning tasks. It distinguishes between different symmetric designs and shows that the reconstruction error of orthonormal symmetric autoencoders can be interpreted via the Eckart‑Young‑Schmidt theorem. Building on this insight, the authors propose an EYS‑based initialization strategy using repeated SVD, and validate its effectiveness through numerical experiments comparing it to conventional deep autoencoders.

By Simone Brivio, Nicola Rares Franco
arXiv Machine Learning
1d ago

Evasion Attacks: How Adversarial Noise Bypasses ML Classifiers

The paper reports a reproducible study of evasion attacks on image and text classifiers. A compact convolutional network on MNIST achieved 98.63% clean accuracy but dropped to 60.20% under FGSM with ε=0.15 and 1.72% with ε=0.30, while PGD reduced accuracy to 32.47% and 0.41%; a bit‑depth‑reduction defense only partially restored performance. In contrast, a DistilBERT model fine‑tuned on the SMS Spam Collection reached 98.75% accuracy and 94.96% F1‑score, yet a sequence of predefined perturbations produced only modest probability shifts and did not flip spam to ham predictions.

By Parker Hummel (Minot State University), Ryne Skabo (Minot State University), Muhammad Abusaqer (Minot State University)
arXiv AI
2d ago

Decision Titan: Test-Time Training for Long-Term Memory in Offline Reinforcement Learning

The paper introduces Decision Titan, a variant of the Decision Transformer that incorporates Test‑Time Training (TTT) layers to store episodic memories in network parameters. It evaluates this architecture on the X‑Maze environment, showing that Decision Titan can learn long‑term dependencies up to 20 times longer than its context window and generalise to sequences 1.7 times longer than the training data. The study also finds that temporal generalisation depends on the choice of time embeddings and that the ability to learn long‑term dependencies hinges on how relevant information is encoded.

By Jude Waide, Robert Lieck
arXiv AI
2d ago

Legal text classification in Korean sexual offense cases: from traditional machine learning to large language models with XAI insights

The paper evaluates legal text classification models for Korean sexual offense cases, comparing traditional machine learning, large language models, and fine‑tuned domain models. Fine‑tuned KLUE‑BERT achieved the highest accuracy of 99.3%, outperforming GPT‑3.5, GPT‑4.0, and other traditional approaches. Explainable AI techniques were used to analyze predictions, revealing linguistic features that influence decisions and highlighting limitations in capturing subtle textual cues, especially in real‑world KICS data.

By Jeongmin Lee
arXiv AI
2d ago

CAVE-Mem: Boundary-Aware Experience Validation for Memory Search

CAVE-Mem is a training‑free framework that enhances memory search for long‑term memory agents by treating experience as a typed intervention operator with conditions on applicability, boundary, and utility. It first retrieves a base answer and then only applies an intervention if the operator matches the current memory substrate, answer contract, evidence boundary, and cross‑fitted utility; otherwise it abstains. Experiments on conversational memory, multi‑hop QA, and long‑document reasoning demonstrate consistent improvements over relevance‑only experience reuse.

By Xinyu Li
arXiv AI
2d ago

MWOP: Modality-aware Width-wise Operation Pruning for Efficient MLLMs

MWOP (Modality-aware Width-wise Operation Pruning) is a method that independently prunes visual‑to‑visual, text‑to‑visual, and text‑to‑text attention paths within each layer of multimodal large language models, and separately selects feed‑forward network channels for visual and textual inputs. It uses a first‑order Taylor criterion to guide pruning, re‑evaluates FFN importance after attention pruning, and applies LoRA‑based recovery training. The approach is paired with path‑sparse Triton attention kernels and compact visual‑side FFN execution to achieve practical acceleration, preserving token sequences while reducing computation. "whyItMatters":"MWOP achieves a 1.6× prefill speedup on LLaVA‑OneVision‑7B while retaining 99.7% performance, and further boosts token‑compression methods to 2.9× and 2.7× speedups, demonstrating its effectiveness across architectures."

By Xudong Wang, Hao Wu, Haozhe Hu, Peiran Yin, Xinghao Chen, Yunpu Ma, Wei Zhang, Xiaoyu Shen
arXiv Computer Vision
2d ago

FocusGraph: Graph-Structured Frame Selection for Embodied Long Video Question Answering

arXiv:2603.04349v2 Announce Type: replace Abstract: Understanding long videos is crucial for embodied intelligent agents, as their performance depends on effectively accumulating and using long-horiz...

By Tatiana Zemskova, Solomon Andryushenko, Ilya Obrubov, Viktoriia Khoruzhaia, Ekaterina Eroshenko, Ekaterina Derevyanka, Dmitry Yudin
arXiv Computer Vision
2d ago

SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation

arXiv:2511.06754v4 Announce Type: replace-cross Abstract: Inspired by how humans reason over discrete objects and their relationships, we explore whether compact object-centric and object-relation re...

By Taisei Hanyu, Nhat Chung, Huy Le, Toan Nguyen, Yuki Ikebe, Anthony Gunderman, Duy Nguyen Ho Minh, Khoa Vo, Tung Kieu, Kashu Yamazaki, Chase Rainwater, Anh Nguyen, Ngan Le
arXiv AI
2d ago

SONIC-O1: A Real-World Benchmark for Evaluating Multimodal Large Language Models on Audio-Video Understanding

SONIC‑O1 is a new benchmark designed to evaluate multimodal large language models on audio‑video understanding. It contains 60 hours of 231 clips across 13 real‑world conversational domains, with 4,958 human‑verified annotations and demographic metadata. The benchmark tests open‑ended summarization, multiple‑choice question answering, and temporally grounded reasoning, revealing performance gaps between model families and across demographic groups.

By Ahmed Y. Radwan, Christos Emmanouilidis, Hina Tabassum, Deval Pandya, Shaina Raza
arXiv AI
2d ago

Ontology-Grounded, Reasoner-Verified Benchmarks for Evaluating LLM Reasoning in Scientific AI

The paper introduces a pipeline that automatically creates ontology‑grounded multiple‑choice question benchmarks for evaluating large language models (LLMs) on logical reasoning tasks in scientific AI. By using OWL 2 ontologies, correct answers are guaranteed by design and distractors are generated and formally verified as incorrect through an OWL reasoner. Experiments on three ontologies—Pizza, PMDco, and DOID—yielded 112, 2,491, and 15,216 MCQs, respectively, with high natural‑language quality and challenging zero‑shot performance for six LLMs.

By Nishtha N. Vaidya, Stephan Grimm, Thomas Hubauer, Thomas A. Runkler
arXiv AI
2d ago

A Comparative Explainability Framework for DeBERTa-v3 in Zero-Shot Medical Abstract Classification

The paper introduces a comparative explainability framework for auditing DeBERTa‑v3 in zero‑shot medical abstract classification. It evaluates five explanation methods—SHAP, LIME, occlusion, Input × Gradient, and Attention × Gradient—using a natural language inference engine on a balanced corpus of 1,000 abstracts per diagnostic category. The study finds that explanatory stability aligns with predictive certainty, identifies three systemic failure mechanisms, and recommends combining multiple explanation methods and quantitative agreement metrics for transformer‑based medical text classifiers.

By Javier Diaz Esteban-Herreros, David Mu\~noz-Valero, Raquel Mart\'inez-Espa\~na, Jose M. Juarez, Juan Moreno-Garcia