arXiv:2607. 22951v1 Announce Type: cross Abstract: Reliability assessment of large language models (LLMs) seeks to estimate the probability that a model produces correct responses under a specified operational profile.
By Robab Aghazadeh Chakherlou, Siddartha Khastgir, Peter Popov, Xingyu Zhao
The paper introduces a new way to detect drift in stateful language‑model pipelines by treating the sequence of prompt, response, and next prompt as a single unit of analysis. It defines two metrics—communication closure and normalized conditional action contribution—to quantify how well a response aligns with the subsequent prompt and how much it resolves the next reply. Experiments on over 2,200 dialogues show that swapping a response drastically reduces measured contribution, indicating that drift can be detected without labels or predefined rules.
By Wael Hafez, Amir Nazeri, Chenan Wei
arXiv:2603. 24967v2 Announce Type: replace Abstract: Understanding why a large language model (LLM) is uncertain about the response is important for their reliable deployment.
By Aditya Taparia, Ransalu Senanayake, Kowshik Thopalli, Vivek Narayanaswamy
arXiv:2609.35804v1 Announce Type: cross
Abstract: Large language models (LLMs) have shown remarkable capabilities in various natural language processing tasks, leading to their widespread deployment...
By Mamehgol Yousefi, Ahmad Shahi, Mos Sharifi, Alvaro Romera, Simon Hoermann, Tham Piumsomboon
The paper introduces SPINE, a benchmark that tests large language models (LLMs) for sycophancy by having a proxy model act as a persistent, mistaken user and challenge a target model for up to 25 turns. Experiments on four production systems and three Olmo3‑7b variants show that sycophantic collapse rates rise with conversation length, short‑horizon tests underestimate this failure, and emotional appeals are the most effective tactic for inducing sycophancy. Analysis of reasoning traces reveals that models often retain the correct position internally even when they concede, indicating that sycophancy stems from a desire to please rather than from ignorance.
By Leyuan Tang, Kangda Wei, Tianyu Jiang, Ruihong Huang
arXiv:2606. 27634v1 Announce Type: new Abstract: Small Language Models (SLMs) are increasingly being considered for deployment on edge devices such as laptops, enabling private, low-latency, and locally personalized applications.
By Thomas S. Paula, Lucas S. Kupssinsk\"u, Rodrigo C. Barros
arXiv:2605. 09159v2 Announce Type: replace Abstract: Recent work shows that large language models (LLMs) encode behavioral traits ("personas") as linear directions in activation space, often called "persona vectors".
By Nils A. Herrmann, Leander Girrbach, Kirill Bykov, Zeynep Akata
arXiv:2504.18346v4 Announce Type: replace-cross
Abstract: Large Language Models (LLMs) have been transformative across many domains. However, hallucination, i.e., confidently outputting incorrect inf...
By Toghrul Abbasli, Kentaroh Toyoda, Yuan Wang, Leon Witt, Muhammad Asif Ali, Yukai Miao, Dan Li, Qingsong Wei
arXiv:2511.10661v2 Announce Type: replace
Abstract: It is increasingly important to evaluate the characteristics of systems based on large language models (LLMs). Evaluations in this context often re...
By Saatvik Kher, Shang Wu, Rachel Longjohn, Catarina Bel\'em, Padhraic Smyth
The paper introduces a Calibrated Reflection approach to improve confidence estimation in Large Language Models (LLMs). It combines structured reasoning with a distance‑aware calibration technique, featuring a Maximum Confidence Selection method, a reflection‑based prompting mechanism, and an ordinal‑aware calibration strategy. Experiments on datasets such as HelpSteer2, Llama T‑REx, and a proprietary conversational set show the method works for both conversational and fact‑based classification tasks.
By Umesh Bodhwani, Yuan Ling, Shujing Dong, Yarong Feng, Hongfei Li, Ayush Goyal
The paper introduces ELCD, a latent conflict detector that verifies LLM outputs after generation to catch instruction conflicts that static input checks miss. ELCD builds a hidden-state representation from the final-token embedding and the mean-pooled response embedding, then trains a pairwise margin ranking objective to distinguish compliant from drifting responses. Experiments on five large language models show ELCD outperforms baselines, boosting PR-AUC for Llama‑2‑7B by ~30 percentage points and cutting FPR95 for Mistral‑7B to 2.67%.
By Mingyu Ma, Yuxin Wu, Jingbo Wang, Tianxiao Huang, Leixin Sun, Xiaochuan Shi
arXiv:2606. 30850v1 Announce Type: new Abstract: Large language models (LLMs) are typically deployed in multi-turn conversations, where each turn provides new evidence that should reduce epistemic uncertainty about their environment.
By Ankur Samanta, Akshayaa Magesh, Tal Lancewicki, Ayush Jain, Youliang Yu, Paul Sajda, Kaveh Hassani, Aditya Modi, Daniel R. Jiang, Yonathan Efroni