The paper introduces Universal Byte-Level Encoding (UBE), a dual‑alphabet tokenizer that routes 3‑4‑byte UTF‑8 characters through UTF‑16 while keeping 1‑2‑byte characters on the UTF‑8 path. This design lowers the worst‑case token‑budget disparity for high‑premium scripts without increasing costs for efficient English spans, and it preserves standard BPE merges and exact decoding. In extensive Unicode audits and multilingual language‑model experiments, UBE matches or improves token‑count efficiency and context usability compared to traditional byte‑pair encoding.
By Hyunsik Kim, Youngmoon Jung
The paper investigates why the train‑validation performance gap widens during fine‑tuning of pretrained models. It proposes a dynamic structural explanation: as training proceeds, updates shift from broadly reusable features to more example‑specific ones, increasing gradient heterogeneity and the gap. Experiments on synthetic ResMLP hierarchies, NLP models (RoBERTa, DeBERTa, Qwen) across six datasets, and vision models (ResNet‑18) confirm that higher reliance on private features correlates with larger accuracy gaps, supporting the proposed account.
By Yuchen Li, Mingyu Du, Zongqi Fan, Ken-Tye Yong, Nguyen H. Tran
The paper presents a theoretical analysis of symmetric autoencoders, a class of deep learning architectures frequently used in machine learning tasks. It distinguishes between different symmetric designs and shows that the reconstruction error of orthonormal symmetric autoencoders can be interpreted via the Eckart‑Young‑Schmidt theorem. Building on this insight, the authors propose an EYS‑based initialization strategy using repeated SVD, and validate its effectiveness through numerical experiments comparing it to conventional deep autoencoders.
By Simone Brivio, Nicola Rares Franco
The paper reports a reproducible study of evasion attacks on image and text classifiers. A compact convolutional network on MNIST achieved 98.63% clean accuracy but dropped to 60.20% under FGSM with ε=0.15 and 1.72% with ε=0.30, while PGD reduced accuracy to 32.47% and 0.41%; a bit‑depth‑reduction defense only partially restored performance. In contrast, a DistilBERT model fine‑tuned on the SMS Spam Collection reached 98.75% accuracy and 94.96% F1‑score, yet a sequence of predefined perturbations produced only modest probability shifts and did not flip spam to ham predictions.
By Parker Hummel (Minot State University), Ryne Skabo (Minot State University), Muhammad Abusaqer (Minot State University)
The paper introduces Decision Titan, a variant of the Decision Transformer that incorporates Test‑Time Training (TTT) layers to store episodic memories in network parameters. It evaluates this architecture on the X‑Maze environment, showing that Decision Titan can learn long‑term dependencies up to 20 times longer than its context window and generalise to sequences 1.7 times longer than the training data. The study also finds that temporal generalisation depends on the choice of time embeddings and that the ability to learn long‑term dependencies hinges on how relevant information is encoded.
By Jude Waide, Robert Lieck
The paper evaluates legal text classification models for Korean sexual offense cases, comparing traditional machine learning, large language models, and fine‑tuned domain models. Fine‑tuned KLUE‑BERT achieved the highest accuracy of 99.3%, outperforming GPT‑3.5, GPT‑4.0, and other traditional approaches. Explainable AI techniques were used to analyze predictions, revealing linguistic features that influence decisions and highlighting limitations in capturing subtle textual cues, especially in real‑world KICS data.
By Jeongmin Lee
CAVE-Mem is a training‑free framework that enhances memory search for long‑term memory agents by treating experience as a typed intervention operator with conditions on applicability, boundary, and utility. It first retrieves a base answer and then only applies an intervention if the operator matches the current memory substrate, answer contract, evidence boundary, and cross‑fitted utility; otherwise it abstains. Experiments on conversational memory, multi‑hop QA, and long‑document reasoning demonstrate consistent improvements over relevance‑only experience reuse.
By Xinyu Li
MWOP (Modality-aware Width-wise Operation Pruning) is a method that independently prunes visual‑to‑visual, text‑to‑visual, and text‑to‑text attention paths within each layer of multimodal large language models, and separately selects feed‑forward network channels for visual and textual inputs. It uses a first‑order Taylor criterion to guide pruning, re‑evaluates FFN importance after attention pruning, and applies LoRA‑based recovery training. The approach is paired with path‑sparse Triton attention kernels and compact visual‑side FFN execution to achieve practical acceleration, preserving token sequences while reducing computation.
"whyItMatters":"MWOP achieves a 1.6× prefill speedup on LLaVA‑OneVision‑7B while retaining 99.7% performance, and further boosts token‑compression methods to 2.9× and 2.7× speedups, demonstrating its effectiveness across architectures."
By Xudong Wang, Hao Wu, Haozhe Hu, Peiran Yin, Xinghao Chen, Yunpu Ma, Wei Zhang, Xiaoyu Shen
arXiv:2610.01921v1 Announce Type: cross
Abstract: Cross-lingual contrastive learning has been a core component of multilingual encoder training, but the ability to explicitly align representations is...
By Lucas Bandarkar, Clark Peng, Ahmed Haj Ahmed, Aditi Khandelwal, Nanyun Peng
arXiv:2602.03006v3 Announce Type: replace
Abstract: Deploying Large Language Models (LLMs) for discriminative workloads is often limited by inference latency, compute, and API costs at scale. Active...
By Ziyang Yu, Liang Zhao
arXiv:2604.07639v2 Announce Type: replace-cross
Abstract: Broadly applicable quantum advantage, particularly in classical data processing and machine learning, has been a fundamental open problem. In...
By Haimeng Zhao, Alexander Zlokapa, Hartmut Neven, Ryan Babbush, John Preskill, Jarrod R. McClean, Hsin-Yuan Huang
arXiv:2609.39055v1 Announce Type: new
Abstract: How can we personalize a shared expert library from a user's pairwise choices? Prior work can realize different reward trade-offs by merging reward-spe...
By Yaling Shen, Tongtong Wu, Siyuan Yan, Gholamreza Haffari
arXiv:2610.01637v1 Announce Type: new
Abstract: In recent decades, artificial intelligence has made significant progress in understanding and interacting with images. One of the important application...
By Cong Phu Nguyen, Huy Tien Nguyen, Tung Le
arXiv:2603.04349v2 Announce Type: replace
Abstract: Understanding long videos is crucial for embodied intelligent agents, as their performance depends on effectively accumulating and using long-horiz...
By Tatiana Zemskova, Solomon Andryushenko, Ilya Obrubov, Viktoriia Khoruzhaia, Ekaterina Eroshenko, Ekaterina Derevyanka, Dmitry Yudin
arXiv:2511.06754v4 Announce Type: replace-cross
Abstract: Inspired by how humans reason over discrete objects and their relationships, we explore whether compact object-centric and object-relation re...
By Taisei Hanyu, Nhat Chung, Huy Le, Toan Nguyen, Yuki Ikebe, Anthony Gunderman, Duy Nguyen Ho Minh, Khoa Vo, Tung Kieu, Kashu Yamazaki, Chase Rainwater, Anh Nguyen, Ngan Le
arXiv:2604.16067v2 Announce Type: replace-cross
Abstract: Fine-tuning pre-trained Vision-Language Models (VLMs) for robotic manipulation introduces a fundamental stability-plasticity dilemma: continu...
By Guransh Singh
arXiv:2605.02035v3 Announce Type: replace-cross
Abstract: Ambiguity resolution is a key challenge in multimodal machine translation (MMT), where models must genuinely leverage visual input to map an...
By Jingheng Pan, Xintong Wang, Longyue Wang, Liang Ding, Weihua Luo, Chris Biemann
SONIC‑O1 is a new benchmark designed to evaluate multimodal large language models on audio‑video understanding. It contains 60 hours of 231 clips across 13 real‑world conversational domains, with 4,958 human‑verified annotations and demographic metadata. The benchmark tests open‑ended summarization, multiple‑choice question answering, and temporally grounded reasoning, revealing performance gaps between model families and across demographic groups.
By Ahmed Y. Radwan, Christos Emmanouilidis, Hina Tabassum, Deval Pandya, Shaina Raza
The paper introduces a pipeline that automatically creates ontology‑grounded multiple‑choice question benchmarks for evaluating large language models (LLMs) on logical reasoning tasks in scientific AI. By using OWL 2 ontologies, correct answers are guaranteed by design and distractors are generated and formally verified as incorrect through an OWL reasoner. Experiments on three ontologies—Pizza, PMDco, and DOID—yielded 112, 2,491, and 15,216 MCQs, respectively, with high natural‑language quality and challenging zero‑shot performance for six LLMs.
By Nishtha N. Vaidya, Stephan Grimm, Thomas Hubauer, Thomas A. Runkler
The paper introduces a comparative explainability framework for auditing DeBERTa‑v3 in zero‑shot medical abstract classification. It evaluates five explanation methods—SHAP, LIME, occlusion, Input × Gradient, and Attention × Gradient—using a natural language inference engine on a balanced corpus of 1,000 abstracts per diagnostic category. The study finds that explanatory stability aligns with predictive certainty, identifies three systemic failure mechanisms, and recommends combining multiple explanation methods and quantitative agreement metrics for transformer‑based medical text classifiers.
By Javier Diaz Esteban-Herreros, David Mu\~noz-Valero, Raquel Mart\'inez-Espa\~na, Jose M. Juarez, Juan Moreno-Garcia