arXiv:2610.07774v1 Announce Type: new
Abstract: Vision-language models (VLMs) face compositional safety risks where harmful intent emerges from the interaction between visual and textual inputs. As m...
By Ziyuan Yang, Wenxuan Ding, Shangbin Feng, Yulia Tsvetkov
arXiv:2610.07847v1 Announce Type: new
Abstract: As LLMs increasingly assist in moral reasoning, omission bias, the tendency to prefer inaction even when equivalent framings reverse substantive outcom...
By Sihyeon Lee, Jihun Song, Chanwoo Kim, Jiwoo Kum, Chanjun Park
arXiv:2610.07936v1 Announce Type: new
Abstract: Systematicity, the probabilistic mapping of form to meaning, permeates language at all levels, and sublexical cues have been shown to govern human pseu...
By Jing Chen, Giulia Loca, Simona Amenta, Marco Marelli
arXiv:2610.08026v1 Announce Type: new
Abstract: In recent years, several methods for detecting when large language models (LLMs) hallucinate have been developed. These methods are often benchmarked w...
By Jorma Valjakka, Juhani Kivim\"aki, Juha Myll\"ari, Jukka K. Nurminen
arXiv:2610.08303v1 Announce Type: new
Abstract: Current evaluation of multilingual Large Language Models (LLMs) rests on an implicit Translation-Isomorphism Assumption (TIA): that semantic structures...
By Shu-Kai Hsieh, Da-Chen Lian
arXiv:2610.08604v1 Announce Type: new
Abstract: Automatic Speech Recognition (ASR) systems often show uneven performance across demographic groups, and errors can be especially difficult to address f...
By Ashley E. Bravo-Bravo, Yuchen Zhang, Haralambos Mouratidis, Ravi Shekhar, Monorama Swain
arXiv:2610.08675v1 Announce Type: new
Abstract: Financial reports repeat values across periods, metrics and accounting lines, allowing an LLM-generated calculation to be numerically correct while cit...
By Chuhong Xu (Sofia University), Bo Su (Indiana University), Ziyao Chen (University of California, San Diego), Ruiyang Xu (Northeastern University), Shimeng Dai (Michigan State University), Xinyu Qiu (Northeastern University)
The study evaluates probabilistic machine‑learning models for forecasting German grid redispatch volumes under data‑delay constraints. Using 48,242 records from 2021‑2024, boosted‑tree LightGBM with rolling calibration achieved the best performance (nWIS 0.7767), outperforming autoregressive and seasonal baselines. Neural models with zero‑censored outputs performed similarly but revealed undercoverage during high‑volume events.
By Faraz Shamim (KIST Medical College and Teaching Hospital, Nepal), Faris Shamim (OTH Regensburg)
The paper introduces Hierarchical Continuous Diffusion Language Models (H-CDLMs), a framework that jointly diffuses multiple token modalities—individual tokens and coarser clusters of token embeddings—to enhance continuous diffusion language models. Applied to the CoBit architecture, the resulting H-CoBit achieves significant empirical gains, improving MAUVE scores and achieving lower generative perplexity on LM1B and OWT, while also outperforming prior continuous diffusion models on GSM8K. The approach generalizes to other continuous generative paradigms, as shown by consistent improvements when applied to the flow matching model FLM.
By Mathias Ollu, Nikos Komodakis
Prefill‑only decision models evaluate every candidate in a menu in a single forward pass, avoiding decoding and reducing cost by one to two orders of magnitude compared to generative language models. The paper demonstrates that when only the candidate menu changes, the model’s post‑intervention accuracy can be predicted solely from the cached first‑pass distribution using a simple estimator that renormalizes and selects the argmax, without any labels or second pass. Across seven model families, ten datasets, and two task types, this menu‑only intervention prediction is within 4.2 points of actual accuracy, and in one family it is exact, whereas a probability‑level variant fails by 21 points, indicating the property resides in ranking rather than calibrated probabilities.
By Ran Li, Lei Chen
The study introduces MedQADE, a German open‑response clinical benchmark with 3,800 question‑answer pairs and physician reference annotations. It evaluates large language models (LLMs) as judges, finding that while some LLMs (e.g., Gemini 3 Flash) achieve physician‑level agreement on correctness, they exhibit self‑bias and low abstention rates. Physicians showed moderate agreement on correctness but limited agreement on difficulty, and their abstention increased with perceived difficulty.
By William Philipp, Finn Fassbender, Daniel Fister, Thorsten Langer, Martje G. Pauly, Rebecca Herzog, Markus A. Hobert, Theresa Paulus, Alexander Baumann, Chi Wang Ip, Lukas L. Goede, Johanna Reimer, Sebastian L\"ons, Ronald B\"ock, Sebastian Fudickar
The paper investigates how vision‑language models acquire and specialize semantic capabilities during fine‑tuning, proposing a trajectory‑based framework that separates acquisition, optima, and specialization phases. It introduces Structured Semantic Routing (SSR) to analyze how supervision representation affects what is learned, showing that SSR improves name‑free attribute‑profile retrieval and that different capabilities peak at different training stages. The study reveals that continued optimization can preserve target‑class retrieval while diminishing transferable semantic knowledge, highlighting a trade‑off between specialization and generalization.
By Suguru Onda, Matthew Bailey, Ryan Farrell
The paper introduces Latent-Action-Guided Video-Language Feature Learning (LAG-VLFL) for recognizing surgical instrument–tissue interactions. By compressing frame-to-frame feature changes into latent actions and predicting next‑frame features, the method aligns video and textual action descriptions without requiring extra spatial or motion annotations. Experiments show that LAG-VLFL improves interaction grounding, temporal‑direction sensitivity, and achieves competitive recognition with faster inference and lower INT4 accuracy loss compared to V‑JEPA2/2.1.
By Jiajun Cheng, Sainan Liu, Subarna Tripathi, Xiaofan Yu, Shan Lin
arXiv:2610.04687v2 Announce Type: replace
Abstract: Answering questions over imperfect tables requires handling errors that can affect the answer. We investigate two challenges for large language mod...
By Baowen Zhang, Wei Fan, Ruman Wang, Hangting Ye
arXiv:2610.06625v2 Announce Type: replace
Abstract: Large language models (LLMs) annotate and scale political text or constructs by generating text tokens. A new class of models, which TypeSafe marke...
By Steven Denney, Matthew DiGiuseppe
arXiv:2610.06729v2 Announce Type: replace
Abstract: Large language models (LLMs) can generate accurate responses, but these are void of diversity. We attempt to address this for the task of creative...
By Zahra Solati Dehkordi, Vasileios Lampos
arXiv:2610.06896v1 Announce Type: new
Abstract: Frontier multimodal large language models (MLLMs) are increasingly positioned as general purpose visual reasoners as part of the quest for artificial g...
By Ross Callaghan, Niannu Gao, Hojjat Azadbakht, Hui Zhang
arXiv:2610.06945v1 Announce Type: new
Abstract: Mechanistic interpretability often relies on the Linear Representation Hypothesis (LRH), which assumes that high-level concepts are encoded as linear d...
By Muhammad Atif Butt, Pawe{\l} Skier\'s, Joost Van De Weijer, Kamil Deja
arXiv:2610.07031v1 Announce Type: new
Abstract: Recent video world models have witnessed the paradigm shift from single-agent to multi-agent involvements, which can reveal more complicated dynamics a...
By Sitian Shen, Jiuming Liu, Mengmeng Liu, Yian Wang, Michael Ying Yang, Francesco Nex, Hao Cheng, Daniele De Martini, Ayush Tewari, Per Ola Kristensson
arXiv:2610.07381v1 Announce Type: new
Abstract: Modeling 3D scene geometry and its evolution over time is essential for autonomous driving and robotics. A common paradigm is to use world models to pr...
By Mehrdad Noori, Guile Wu, Sam Hosseini, Dongfeng Bai