arXiv:2608.28840v1 Announce Type: new
Abstract: Independently trained neural networks tend to encode the same data with similar latent geometries. These latent geometries are not directly compatible,...
By Cameron Ryan, Vivek Sivaraman Narayanaswamy, Kowshik Thopalli, Shusen Liu
arXiv:2606.16206v2 Announce Type: replace
Abstract: Large language models are increasingly proposed as educational tutors, yet stronger task-solving ability does not necessarily imply stronger learni...
By Junyi Yao, Zihao Zheng, Baichuan Li
arXiv:2604.12335v2 Announce Type: replace-cross
Abstract: Training multimodal large language models (MLLMs) for video understanding requires large-scale annotated data spanning diverse tasks such as...
By Tanzila Rahman, Renjie Liao, Leonid Sigal
arXiv:2608.30975v1 Announce Type: cross
Abstract: Cardiac magnetic resonance imaging (CMR) produces rich sequential data such as temporal cine videos and spatial LGE/mapping stacks, yet most deep lea...
By Athira J. Jacob, Puneet Sharma, Dorin Comaniciu, Daniel Rueckert
arXiv:2605.28740v2 Announce Type: replace-cross
Abstract: As large language models are increasingly deployed for clinical text, ensuring they can reliably signal their own uncertainty becomes critica...
By Bushi Xiao, Sarvesh Soni, Daisy Zhe Wang
ManGo is an unsupervised framework for manga visual question answering that actively selects panels, extracts concise clues, and decides when to stop, creating a compact evidence sketch before answering. It introduces Active Narrative Sketching (ANS) and optimizes its behavior using group-relative policy training with two rewards: answer preference from listwise self-ranking and path consistency from stable ordered panel trajectories. Experiments on standard manga understanding benchmarks demonstrate that ManGo achieves state‑of‑the‑art performance across different settings.
By Hao Qiu, Junyan Wang, Zheyuan Liu, Lei Fan, Hong Jia, Lianbo Guo, Zhulin Tao
arXiv:2601.02933v4 Announce Type: replace
Abstract: Human evaluation is the gold standard for multilingual NLP, but is often skipped in practice and substituted with automatic metrics because it is n...
By Vil\'em Zouhar, Tom Kocmi
General-purpose NLP embedding models perform well on linguistic tasks, but their ability to capture symbolic ontological structure remains unclear. We introduce AVA, a systematic framework for evaluat...
We present Qwen-Drive-1.0, an initial step towards a vision-language foundation model for autonomous driving. Qwen-Drive-1.0 retains the architecture of the pretrained vision-language model (VLM) and...
Artificial intelligence (AI) and natural language processing (NLP) are increasingly used to extract, integrate, and interpret biomedical knowledge relevant to cancer genomics, yet their translation in...
Building language technologies and conducting NLP research for low-resource languages---particularly when led by native speakers or involving participatory research practices---are often framed as mea...
Large language models (LLMs) are built from structured high-dimensional objects such as token representations, weights, adaptation updates, caches, and activations, whose multilinear structure is unde...
We present an overview of the ImageEval 2026 shared task on culturally grounded Arabic multimodal evaluation. It includes two tasks: (i) AynVQA, covering spoken visual question answering and image-gro...
Indirect speech acts (ISAs) require pragmatic reasoning over context, as directive intent can- not be inferred from surface form alone. Prior text-based studies and existing multimodal benchmarks larg...
Long-context LLM applications such as document summarization and multi-turn agents require generation from prefixes spanning tens of thousands of tokens, making decoding latency a major bottleneck. Sp...
The paper introduces Short‑Films 20K (SF20K), a large publicly available movie dataset comprising 20,143 amateur films totaling 3,582 hours, with an average length of 12 minutes per film. Accompanying the dataset is SF20K‑Test, a manual open‑ended question‑answering benchmark featuring 95 movies and 979 question‑answer pairs. Analysis of the benchmark shows limited data leakage, highlights the necessity of long‑term reasoning, and demonstrates that instruction tuning on the large‑scale dataset significantly boosts vision‑language model performance.
By Ridouane Ghermi, Xi Wang, Vicky Kalogeiton, Ivan Laptev
The paper introduces Active Testing, a framework that selects the most informative test samples for annotation in NLP, aiming to reduce human effort while accurately estimating model performance. Experiments across 18 datasets and 4 embedding strategies show up to 95% annotation savings with less than 1% loss in performance estimation accuracy. The authors also propose an adaptive stopping criterion to determine the optimal number of samples without a predefined budget.
By Antonio Purificato, Maria Sofia Bucarelli, Andrea Bacciu, Fabrizio Silvestri, Amin Mantrach
The paper explores using language models to restore missing text in damaged ancient manuscripts caused by physical gaps. It tests various scenarios, model architectures, and decoding strategies to handle tokenization mismatches and lacuna length awareness. Results show that while full automation is not yet possible, these tools can effectively aid paleographers, with performance varying by document section and missing text length.
By Shibingfeng Zhang, Edoardo Caraffa, Annafelicia Zuffrano, Maddalena Modesti, Giovanni Colavizza
The paper proposes the Effectiveness–Losslessness Framework to guide tokenization in domains beyond language, using predictive codelength as a criterion. It introduces two boundaries: the Fact–Token Boundary, where observable structure should be encoded into tokens, and the Token–State Boundary, where context‑dependent relations should remain for model state rather than being pre‑tokenized. Experiments on symbolic music show that making musical time explicit and applying tonal‑frame canonicalization improve predictive performance, while fixed pitch coordinates and reversible BPE can increase predictive code length, indicating that carrier compaction alone does not guarantee better predictions.
By Yi Wang
The paper introduces a new protocol, Mistake Detection Video Question Answering (MD‑VQA), to evaluate whether models can determine if a step in a video follows its description, covering both seen and unseen actions. It proposes a post‑training approach for video‑language models that uses a reward function to highlight discrepancies between instructions and video content. Experiments show this method surpasses zero‑shot, fine‑tuned, and other post‑training baselines, especially on unseen procedures, improving performance by up to 11.6% on EP‑VQA.
By Federico Spurio, Olga Zatsarynna, Lars Doorenbos, Emad Bahrami, Gianpiero Francesca, Juergen Gall