arXiv:2604.02691v2 Announce Type: replace
Abstract: Deep learning based semantic communication has achieved significant progress in wireless image transmission, but most existing schemes rely on fixe...
By Haowen Wan, Qianqian Yang
arXiv:2607.09042v2 Announce Type: replace
Abstract: Reinforcement learning is increasingly used to fine-tune vision-language-action (VLA) models, but robot interaction is expensive and learning becom...
By Iris Xu, Sunshine Jiang, John Marangola, Pulkit Agrawal, Zhang-Wei Hong
arXiv:2609.34457v2 Announce Type: replace
Abstract: Transformer-based models are widely used for reasoning, coding, and multimodal agentic tasks. To provide formal assurance of desirable behaviors, s...
By Hai Duong, Thanh Le, ThanhVu Nguyen
arXiv:2610.06940v1 Announce Type: new
Abstract: Continual adaptation of language models can change their output distribution on prompts learned earlier, while retaining every old prompt-answer pair m...
By Huan Li, Zhe Cao, Qinlei Xie, Fushun Cui, Xuechen Liang
arXiv:2610.07186v1 Announce Type: new
Abstract: Large language models make claims about themselves that are both consequential and increasingly difficult to verify from behavior alone. How can we dis...
By David I. Atkinson, Dillon Plunkett, David Bau
arXiv:2610.07563v1 Announce Type: new
Abstract: Large language models (LLMs) have the potential to meet a key goal in economics: a quantitative model of household decision making, across a variety of...
By Jin Huang, Diego Ferreras Garrucho, Yutong Xie, Walter M. Yuan, Qiaozhu Mei, Chen Lian, Jonathon Hazell
arXiv:2610.07700v1 Announce Type: new
Abstract: We study the robustness of keystroke dynamics for detecting large language model (LLM)-assisted writing. We introduce a Vietnamese keystroke dataset ca...
By Thanh Dong, An Ngo, Minh Dau, Rajesh Kumar
arXiv:2610.07774v1 Announce Type: new
Abstract: Vision-language models (VLMs) face compositional safety risks where harmful intent emerges from the interaction between visual and textual inputs. As m...
By Ziyuan Yang, Wenxuan Ding, Shangbin Feng, Yulia Tsvetkov
arXiv:2610.07847v1 Announce Type: new
Abstract: As LLMs increasingly assist in moral reasoning, omission bias, the tendency to prefer inaction even when equivalent framings reverse substantive outcom...
By Sihyeon Lee, Jihun Song, Chanwoo Kim, Jiwoo Kum, Chanjun Park
arXiv:2610.07936v1 Announce Type: new
Abstract: Systematicity, the probabilistic mapping of form to meaning, permeates language at all levels, and sublexical cues have been shown to govern human pseu...
By Jing Chen, Giulia Loca, Simona Amenta, Marco Marelli
arXiv:2610.08026v1 Announce Type: new
Abstract: In recent years, several methods for detecting when large language models (LLMs) hallucinate have been developed. These methods are often benchmarked w...
By Jorma Valjakka, Juhani Kivim\"aki, Juha Myll\"ari, Jukka K. Nurminen
arXiv:2610.08303v1 Announce Type: new
Abstract: Current evaluation of multilingual Large Language Models (LLMs) rests on an implicit Translation-Isomorphism Assumption (TIA): that semantic structures...
By Shu-Kai Hsieh, Da-Chen Lian
arXiv:2610.08604v1 Announce Type: new
Abstract: Automatic Speech Recognition (ASR) systems often show uneven performance across demographic groups, and errors can be especially difficult to address f...
By Ashley E. Bravo-Bravo, Yuchen Zhang, Haralambos Mouratidis, Ravi Shekhar, Monorama Swain
arXiv:2610.08675v1 Announce Type: new
Abstract: Financial reports repeat values across periods, metrics and accounting lines, allowing an LLM-generated calculation to be numerically correct while cit...
By Chuhong Xu (Sofia University), Bo Su (Indiana University), Ziyao Chen (University of California, San Diego), Ruiyang Xu (Northeastern University), Shimeng Dai (Michigan State University), Xinyu Qiu (Northeastern University)
The study evaluates probabilistic machine‑learning models for forecasting German grid redispatch volumes under data‑delay constraints. Using 48,242 records from 2021‑2024, boosted‑tree LightGBM with rolling calibration achieved the best performance (nWIS 0.7767), outperforming autoregressive and seasonal baselines. Neural models with zero‑censored outputs performed similarly but revealed undercoverage during high‑volume events.
By Faraz Shamim (KIST Medical College and Teaching Hospital, Nepal), Faris Shamim (OTH Regensburg)
The paper introduces Hierarchical Continuous Diffusion Language Models (H-CDLMs), a framework that jointly diffuses multiple token modalities—individual tokens and coarser clusters of token embeddings—to enhance continuous diffusion language models. Applied to the CoBit architecture, the resulting H-CoBit achieves significant empirical gains, improving MAUVE scores and achieving lower generative perplexity on LM1B and OWT, while also outperforming prior continuous diffusion models on GSM8K. The approach generalizes to other continuous generative paradigms, as shown by consistent improvements when applied to the flow matching model FLM.
By Mathias Ollu, Nikos Komodakis
Prefill‑only decision models evaluate every candidate in a menu in a single forward pass, avoiding decoding and reducing cost by one to two orders of magnitude compared to generative language models. The paper demonstrates that when only the candidate menu changes, the model’s post‑intervention accuracy can be predicted solely from the cached first‑pass distribution using a simple estimator that renormalizes and selects the argmax, without any labels or second pass. Across seven model families, ten datasets, and two task types, this menu‑only intervention prediction is within 4.2 points of actual accuracy, and in one family it is exact, whereas a probability‑level variant fails by 21 points, indicating the property resides in ranking rather than calibrated probabilities.
By Ran Li, Lei Chen
The study introduces MedQADE, a German open‑response clinical benchmark with 3,800 question‑answer pairs and physician reference annotations. It evaluates large language models (LLMs) as judges, finding that while some LLMs (e.g., Gemini 3 Flash) achieve physician‑level agreement on correctness, they exhibit self‑bias and low abstention rates. Physicians showed moderate agreement on correctness but limited agreement on difficulty, and their abstention increased with perceived difficulty.
By William Philipp, Finn Fassbender, Daniel Fister, Thorsten Langer, Martje G. Pauly, Rebecca Herzog, Markus A. Hobert, Theresa Paulus, Alexander Baumann, Chi Wang Ip, Lukas L. Goede, Johanna Reimer, Sebastian L\"ons, Ronald B\"ock, Sebastian Fudickar
The paper investigates how vision‑language models acquire and specialize semantic capabilities during fine‑tuning, proposing a trajectory‑based framework that separates acquisition, optima, and specialization phases. It introduces Structured Semantic Routing (SSR) to analyze how supervision representation affects what is learned, showing that SSR improves name‑free attribute‑profile retrieval and that different capabilities peak at different training stages. The study reveals that continued optimization can preserve target‑class retrieval while diminishing transferable semantic knowledge, highlighting a trade‑off between specialization and generalization.
By Suguru Onda, Matthew Bailey, Ryan Farrell
The paper introduces Latent-Action-Guided Video-Language Feature Learning (LAG-VLFL) for recognizing surgical instrument–tissue interactions. By compressing frame-to-frame feature changes into latent actions and predicting next‑frame features, the method aligns video and textual action descriptions without requiring extra spatial or motion annotations. Experiments show that LAG-VLFL improves interaction grounding, temporal‑direction sensitivity, and achieves competitive recognition with faster inference and lower INT4 accuracy loss compared to V‑JEPA2/2.1.
By Jiajun Cheng, Sainan Liu, Subarna Tripathi, Xiaofan Yu, Shan Lin