The paper questions whether large language models (LLMs) truly introspect by critiquing recent studies that claim they can detect and report their internal states. It proposes two necessary conditions for genuine introspection: privileged access to internal representations and second‑order computation that distinguishes from first‑order task performance. Re‑examining two existing paradigms, the authors find that apparent introspective abilities can be explained by input‑based classifiers or generic anomaly detection, concluding that current evidence does not support metacognitive monitoring in LLMs.
By Shashwat Singh, Tal Linzen, Shauli Ravfogel
arXiv:2607. 23379v1 Announce Type: cross Abstract: Activation Oracles (AOs) are language models trained to answer natural-language questions about another model's internal activations.
By Tobias Bersia, Tatiana Gaintseva
arXiv:2609.22119v1 Announce Type: cross
Abstract: Evaluation awareness poses an unprecedented threat to model evaluation, but the mechanisms by which models detect it remain unknown. This study focus...
By Navraj Singh, Maheep Chaudhary
arXiv:2608. 05624v1 Announce Type: new Abstract: Sycophantic responses are becoming pervasive in large language models (LLMs), and prior work has pointed out that some of them could be harmful.
By Bohan Jiang, Dawei Li, Yasin Silva, Huan Liu
arXiv:2609.31603v1 Announce Type: cross
Abstract: Large language models (LLMs) implicitly infer attributes of their users and adapt their behavior accordingly, yet these beliefs remain difficult to i...
By Ali Holmov, Yiran Huang, Kirill Bykov, Zeynep Akata
arXiv:2608.02657v2 Announce Type: replace-cross
Abstract: Agentic LLMs are vulnerable to indirect prompt injection (IPI) attacks, e.g., malicious side-tasks hidden in external tool results. While man...
By Jianshuo Dong, Yiming Liu, Maosen Zhang, Nan Deng, Peng Xu, Xiaoping Zhang, Tianwei Zhang, Jie Zhang, Han Qiu