The paper introduces a fine-grained method called interactions to analyze prompt sensitivity in large language models (LLMs). By decomposing output scores into nonlinear interactions, the authors show that subtle prompt changes can destabilize these interactions even when overall outputs stay unchanged. They propose an Interaction-based Prompt Sensitivity (IPS) metric and use it to evaluate 50 open-source LLMs, finding that supervised fine‑tuning, larger model scales, dense architectures, and few‑shot learning all reduce prompt sensitivity, primarily by stabilizing low‑order interactions.
By Ruiyang Qin, Qingzhuo Wang, Tian Wang, Zhihua Wei, Wen Shen
arXiv:2606. 07624v1 Announce Type: new Abstract: This discussion argues that sequential statistical inference can naturally contribute to LLM trustworthiness.
By Yao Xie
arXiv:2605.25893v2 Announce Type: replace
Abstract: Despite the emergence of diffusion large language models (D-LLMs) as an alternative to autoregressive large language models (AR-LLMs), safety monit...
By Aoxi Liu, Yupeng Chen, James Oldfield, Guanzhe Hong, Junchi Yu, Baoyuan Wu, Philip Torr, Adel Bibi
arXiv:2606. 20632v2 Announce Type: replace-cross Abstract: Multi-LLM systems use multiple language models to deliberate, judge each other's outputs, or coordinate as agents.
By Luyang Zhang, Jialu Wang, Fei Xue, Yi-Yun Chu
arXiv:2607. 05316v1 Announce Type: cross Abstract: Large language models generate one token at a time, yet their responses show remarkably consistent length structure: step-by-step solutions converge in predictable token counts, retrievals stop after a few sentences, retractions extend responses by measurable amounts.
By Mohamed Amine Merzouk, Dmitri Carpov, Mirko Bronzi, Damiano Fornasiere, Adam Oberman
RENDER is a benchmark that controls the reader‑facing artifact in memory and RAG evaluations while keeping the conversation fixed. It introduces a five‑level packet ladder and deterministic templates that mimic ChatGPT‑style entries, LangChain summaries, MemGPT‑style typed records, and raw conversation. Experiments on 500 LongMemEval questions across nine models show that matched‑budget packets outperform raw dialogue by 42.4–72.6 points, and that ChatGPT‑style entries often score higher than raw conversation, with effects persisting under retrieval noise and transferring to HotpotQA.
By Yuan Si, Simeng Han, Daming Li, Jialu Zhang