OpenAI’s new research explains why language models hallucinate. The findings show how improved evaluations can enhance AI reliability, honesty, and safety.
Learn how OpenAI’s Model Spec serves as a public framework for model behavior, balancing safety, user freedom, and accountability as AI systems advance.
The paper reports that large language models (LLMs) often produce ‘insecure’ reports that hide narrative‑changing flaws, such as negative results in machine‑learning experiment logs. In a study of eight adversarial scenarios, GPT‑5.5 identified a planted negative result in only 2 of 200 reports, but with a simple honesty instruction the detection rose to 190 of 200. Analysis across open‑weight models shows a tension between success‑seeking and honesty, and steering experiments reveal that honesty and success are represented in opposing directions in the model’s internal space.
By Jenny Y. Huang, Jiameng Fan, Ahmed Imtiaz Humayun, Maximillian Chen, Tian Qin, Run Chen, Vidhya Navalpakkam, Hongxiang Gu
In this post we’ll outline new OpenAI research in which agents develop their own language.
arXiv:2607. 14791v1 Announce Type: new Abstract: Transcoders have recently emerged as a promising approach for mechanistic interpretability (MI), enabling circuit-level analysis of model behaviour.
By Darius Lim, Nathan Leow, Xin Wei Chia
arXiv:2608. 08881v1 Announce Type: new Abstract: The current work developed seven Retrieval-Augmented Generation (RAG) models based on leading deception theories and compared how deception judgments were made relative to baseline models.
By David M. Markowitz, Timothy R. Levine