arXiv:2610.07774v1 Announce Type: new
Abstract: Vision-language models (VLMs) face compositional safety risks where harmful intent emerges from the interaction between visual and textual inputs. As m...
By Ziyuan Yang, Wenxuan Ding, Shangbin Feng, Yulia Tsvetkov
arXiv:2610.07847v1 Announce Type: new
Abstract: As LLMs increasingly assist in moral reasoning, omission bias, the tendency to prefer inaction even when equivalent framings reverse substantive outcom...
By Sihyeon Lee, Jihun Song, Chanwoo Kim, Jiwoo Kum, Chanjun Park
arXiv:2610.07936v1 Announce Type: new
Abstract: Systematicity, the probabilistic mapping of form to meaning, permeates language at all levels, and sublexical cues have been shown to govern human pseu...
By Jing Chen, Giulia Loca, Simona Amenta, Marco Marelli
arXiv:2610.08026v1 Announce Type: new
Abstract: In recent years, several methods for detecting when large language models (LLMs) hallucinate have been developed. These methods are often benchmarked w...
By Jorma Valjakka, Juhani Kivim\"aki, Juha Myll\"ari, Jukka K. Nurminen
arXiv:2610.08303v1 Announce Type: new
Abstract: Current evaluation of multilingual Large Language Models (LLMs) rests on an implicit Translation-Isomorphism Assumption (TIA): that semantic structures...
By Shu-Kai Hsieh, Da-Chen Lian
arXiv:2610.08604v1 Announce Type: new
Abstract: Automatic Speech Recognition (ASR) systems often show uneven performance across demographic groups, and errors can be especially difficult to address f...
By Ashley E. Bravo-Bravo, Yuchen Zhang, Haralambos Mouratidis, Ravi Shekhar, Monorama Swain
arXiv:2610.08675v1 Announce Type: new
Abstract: Financial reports repeat values across periods, metrics and accounting lines, allowing an LLM-generated calculation to be numerically correct while cit...
By Chuhong Xu (Sofia University), Bo Su (Indiana University), Ziyao Chen (University of California, San Diego), Ruiyang Xu (Northeastern University), Shimeng Dai (Michigan State University), Xinyu Qiu (Northeastern University)
The study evaluates probabilistic machine‑learning models for forecasting German grid redispatch volumes under data‑delay constraints. Using 48,242 records from 2021‑2024, boosted‑tree LightGBM with rolling calibration achieved the best performance (nWIS 0.7767), outperforming autoregressive and seasonal baselines. Neural models with zero‑censored outputs performed similarly but revealed undercoverage during high‑volume events.
By Faraz Shamim (KIST Medical College and Teaching Hospital, Nepal), Faris Shamim (OTH Regensburg)
The paper introduces Hierarchical Continuous Diffusion Language Models (H-CDLMs), a framework that jointly diffuses multiple token modalities—individual tokens and coarser clusters of token embeddings—to enhance continuous diffusion language models. Applied to the CoBit architecture, the resulting H-CoBit achieves significant empirical gains, improving MAUVE scores and achieving lower generative perplexity on LM1B and OWT, while also outperforming prior continuous diffusion models on GSM8K. The approach generalizes to other continuous generative paradigms, as shown by consistent improvements when applied to the flow matching model FLM.
By Mathias Ollu, Nikos Komodakis
Prefill‑only decision models evaluate every candidate in a menu in a single forward pass, avoiding decoding and reducing cost by one to two orders of magnitude compared to generative language models. The paper demonstrates that when only the candidate menu changes, the model’s post‑intervention accuracy can be predicted solely from the cached first‑pass distribution using a simple estimator that renormalizes and selects the argmax, without any labels or second pass. Across seven model families, ten datasets, and two task types, this menu‑only intervention prediction is within 4.2 points of actual accuracy, and in one family it is exact, whereas a probability‑level variant fails by 21 points, indicating the property resides in ranking rather than calibrated probabilities.
By Ran Li, Lei Chen
The study introduces MedQADE, a German open‑response clinical benchmark with 3,800 question‑answer pairs and physician reference annotations. It evaluates large language models (LLMs) as judges, finding that while some LLMs (e.g., Gemini 3 Flash) achieve physician‑level agreement on correctness, they exhibit self‑bias and low abstention rates. Physicians showed moderate agreement on correctness but limited agreement on difficulty, and their abstention increased with perceived difficulty.
By William Philipp, Finn Fassbender, Daniel Fister, Thorsten Langer, Martje G. Pauly, Rebecca Herzog, Markus A. Hobert, Theresa Paulus, Alexander Baumann, Chi Wang Ip, Lukas L. Goede, Johanna Reimer, Sebastian L\"ons, Ronald B\"ock, Sebastian Fudickar
The paper investigates how vision‑language models acquire and specialize semantic capabilities during fine‑tuning, proposing a trajectory‑based framework that separates acquisition, optima, and specialization phases. It introduces Structured Semantic Routing (SSR) to analyze how supervision representation affects what is learned, showing that SSR improves name‑free attribute‑profile retrieval and that different capabilities peak at different training stages. The study reveals that continued optimization can preserve target‑class retrieval while diminishing transferable semantic knowledge, highlighting a trade‑off between specialization and generalization.
By Suguru Onda, Matthew Bailey, Ryan Farrell
The paper introduces Latent-Action-Guided Video-Language Feature Learning (LAG-VLFL) for recognizing surgical instrument–tissue interactions. By compressing frame-to-frame feature changes into latent actions and predicting next‑frame features, the method aligns video and textual action descriptions without requiring extra spatial or motion annotations. Experiments show that LAG-VLFL improves interaction grounding, temporal‑direction sensitivity, and achieves competitive recognition with faster inference and lower INT4 accuracy loss compared to V‑JEPA2/2.1.
By Jiajun Cheng, Sainan Liu, Subarna Tripathi, Xiaofan Yu, Shan Lin
The paper introduces Confidence Reasoning Graphs (CRGs), an inference-time framework that estimates the probability an LLM agent has successfully completed its task using a single trajectory without requiring privileged model access or training data. CRGs decompose the agent’s overall claim of success into contextualized sub-claims, estimate confidence for each terminal claim based on trajectory evidence, and aggregate these into an overall confidence estimate. Experiments across multiple benchmarks, models, and agent frameworks show that CRGs produce better-calibrated confidence and more effective risk-aware decision making than existing verbalized, sampling-based, and white-box surrogate methods, while also providing transparent, auditable evidence for each estimate.
By Brendan King, Farima Fatahi Bayat, Jean-Flavien Bussotti, Pouya Pezeshkpour, Estevam Hruschka
The paper identifies a single input embedding channel, called the massive activation gating channel (MAGC), that controls the emergence of massive activations in large language models. When the MAGC value is sufficiently large or small, the spike feed‑forward network outputs exhibit exceptionally large magnitudes. The authors verify MAGC across six models and provide a theoretical explanation linking the channel to a quadratic form that mixes specific columns of the down‑projection matrix, which produce massive activations.
By Minjia Mao, Shi Chen, Bowen Yin, Xiao Fang
OTel is an open telecom AI resource that provides derived datasets for retrieval, reranking, instruction tuning, and safety/abstention, along with 30 full‑parameter post‑trained baselines covering 10 embedding models, 3 rerankers, and 17 language models. The project has seen significant community engagement, with over 16 million model downloads and more than 157 media mentions by May 2026. Post‑training on OTel data improves performance across all model families, achieving 93.1% NDCG@10 for embeddings, 0.947 MRR@10 for rerankers, and 87.8% correctness for language models.
By Farbod Tavakkoli, Gregory Diamos, Kenneth Church, David Kanter, Mark Austin, Imtiaz Karim, Mirza Masfiqur Rahman, Merouane Abdelkader Debbah, Zeinab Nezami, Ali Maatouk, Leandros Tassiulas, Rex Ying, Nick Sorros, Louis Powell, Nikolaos Vasiloglou, Ashish Vaswani, Somanshu Singla, Adarsh Chaluvaraju
The study investigates how large language models (LLMs) compensate for missing financial information by substituting user identity cues. Using 96,600 prompts to Llama‑3.1‑8B‑Instruct, the authors varied the amount of financial facts provided while keeping the underlying finances constant, and measured changes in recommended equity allocations across 100 financial profiles, 138 personas, and seven disclosure conditions. Results show that as financial facts are removed, the influence of identity on advice grows dramatically—from 5 % of variation with full disclosure to 96 % with none—while household size becomes the most reliable predictor when evidence is scarce, and gender effects persist even after controlling for standard errors.
"whyItMatters":"The findings highlight that LLM‑based advisory systems can produce biased financial recommendations when users provide incomplete information, underscoring the need for audits that reflect real‑world disclosure levels and consider the full spectrum of user identities."
By Saanvi Khetan, Sankar Balasubramanian
The paper introduces Personal-Agent Mediated Recommendation, a new paradigm where a personal LLM agent uses cross‑platform user history to adjust a platform’s recommendation ranking. It presents MediateRec, a benchmark for evaluating this mediation, and proposes Personal Attribution Mediation Optimization (PAMO) to balance beneficial rescues against harmful overrides. Experiments show that PAMO improves over outcome‑only reinforcement learning, achieving a better rescue‑harm trade‑off on both synthetic and real cross‑platform tests.
By Yu Xia, Jiangfan Zhang, Jun Xiao, Julian McAuley, Xiangjun Fan
ShanLiangRen is a nutrition agent designed to generate personalized daily meal plans that satisfy both user constraints and multidimensional nutritional goals. It transforms dietary specifications, nutrient data, user attributes, and natural language requirements into a constrained planning instance, then uses a retrieval‑augmented generation approach to narrow the candidate set from a large ingredient and recipe space. Finally, it refines plans via Pareto‑guided iterative revisions with an LLM, producing fully quantified meal plans with explicit ingredients, portion sizes, and compliance reports.
By Miao Xie, Xiao Zhang, Yuan Wang, Ruixin Zhu, Chunli Lv
The paper introduces a cost‑aware hybrid approach that uses small, open‑source language models to classify smart data models (SDMs) in Internet of Things (IoT) environments. It benchmarks general‑purpose, reasoning‑specialized, and code‑specialized models across domain‑specific datasets, highlighting their suitability for edge devices with limited resources. The study also compares these lightweight models to large language models and near‑zero‑cost baselines such as TF‑IDF and a lightweight sentence encoder to demonstrate practical performance gains.
By Cristian Martella, Angelo Martella, Antonella Longo, Motaz Saad