arXiv:2609.35860v1 Announce Type: cross
Abstract: Sampling based consistency is widely used for hallucination detection, yet aggregate performance can conceal systematic differences in which errors a...
By Pranav Darshan, Pranav A, Sravan Karthick T, Minal Moharir, Ivan P. Yamshchikov
The asymmetry between language production and perception has been well-documented in psycholinguistics. Whether large language models (LLMs) exhibit a functionally analogous distinction remains an open question, particularly given that LLMs rely on the same underlying mechanism (next-token prediction) for both input and output processing.
arXiv:2606. 30449v1 Announce Type: new Abstract: Probes on model internals could help monitor agentic systems if they identify harmful text or tool actions before those actions are generated.
By Max Fomin, Elad David, Amit LeVi
arXiv:2609. 31181v1 Announce Type: new Abstract: Black-box model identification works by scoring a model's response to natural-language prompts.
By Nicol\'as Vera Z\'u\~niga
arXiv:2608. 04021v1 Announce Type: cross Abstract: Cloze-style probes that vary how often a target token appears implicitly assume that more copies of a target affect prediction the same way regardless of where the readout slot sits.
By Han-yu Wang
arXiv:2609.07474v2 Announce Type: replace
Abstract: Language models compute over tokens: language is their input, their output, and increasingly their internal representation. Whether language should...
By Peng Xie, Amr Alanwar
The study examines how different editorial framings in prompts influence large language models’ statistical analysis reports. Using a 4×4 factorial design, researchers found that certain framings—particularly brutally critical prompts on genuine effects and significance-seeking prompts on underpowered nulls—led to factual misrepresentations. Tone shifts were more widespread, with critical framing inducing defensive language across all data patterns, while a confound in the data largely prevented both factual and tonal distortions.
By Paras Balani, Subhrakanta Panda
Warning: This paper studies stereotypes and biases, and contains potentially disturbing examples, used for illustration purposes only. Our findings should not be interpreted as an argument against alignment.
arXiv:2602. 06941v2 Announce Type: replace-cross Abstract: Large language models can recover mid-generation from task-misaligned activation steering, producing explicit verbal restarts (e.
By Alex McKenzie, Keenan Pepper, Stijn Servaes, Martin Leitgab, Murat Cubuktepe, Mike Vaiana, Diogo de Lucena, Judd Rosenblatt, Michael S. A. Graziano
arXiv:2604. 19139v3 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) continue to evolve through alignment techniques such as Reinforcement Learning from Human Feedback (RLHF) and Constitutional AI, a growing and increasingly conspicuous phenomenon has emerged: the proliferation of verbal tics--repetitive, formulaic linguistic patterns that pervade model outputs.
By Shuai Wu, Xue Li, Yanna Feng, Yufang Li, Zhijun Wang, Ran Wang
arXiv:2606. 07889v1 Announce Type: cross Abstract: LLM-based coding agents sometimes acknowledge a problem in their own reasoning and then proceed anyway.
By Marut Pandya, Kasey Zhang, Baiqing Lyu
arXiv:2606. 12747v1 Announce Type: new Abstract: Safety-relevant studies of language models, including alignment and jailbreaking evaluations and AI control protocols, often rely on prefilling model outputs.
By Andy Wang, Parv Mahajan, David Demitri Africa, Alexandra Souly, Jordan Taylor, Robert Kirk