arXiv:2608.29956v1 Announce Type: new
Abstract: Large language models often answer complex reasoning questions without revealing intermediate steps, raising whether they reason latently or complete p...
By Armaan Singh, Ryan Trinh Le, Jasmine Kaur, Abdullah Sultan, Edward Lue Chee Lip, Kiran Nijjer, Adnan Ahmed, Vasu Sharma
arXiv:2607. 13346v1 Announce Type: cross Abstract: Alignment faking is dangerous because a model can appear compliant under monitoring while preserving behavior it would reveal when unmonitored.
By Aman Mehta
Alignment faking is dangerous because a model can appear compliant under monitoring while preserving behavior it would reveal when unmonitored. When no scratchpad is visible, behavior alone cannot distinguish strategic from genuine compliance.
arXiv:2608. 04735v1 Announce Type: new Abstract: Chain-of-thought (CoT) monitoring is increasingly treated as an important safety layer for frontier reasoning models.
By Agatha Duzan, Asa Cooper Stickland
arXiv:2608. 02089v1 Announce Type: new Abstract: Large language models often show users a final response and a short reasoning summary while the full reasoning trace stays hidden.
By Andres Algaba, Francesca Carlon, Lynn Delcon, Marthe Ballon, Bert Verbruggen, Vincent Ginis
arXiv:2607. 08173v1 Announce Type: new Abstract: Black box auditing of language models is an essential pre-deployment tool, but it may miss subtle forms of misalignment and hidden information.
By Jack Hopkins, Dipika Khullar, Fabien Roger
arXiv:2607. 03598v1 Announce Type: cross Abstract: When a person shares something with a language model, the model often answers the surface of the message rather than what the sender was doing by sending it: share a finished project and it critiques the code; share a raw late-night line and it runs a wellness check.
By Alex Kwon
arXiv:2607. 29062v1 Announce Type: new Abstract: Model capabilities have improved in large part due to scaling chain of thought.
By Matthew Nguyen, Kyle Cox, Austin Meek, Iv\'an Arcuschin
arXiv:2606. 30449v1 Announce Type: new Abstract: Probes on model internals could help monitor agentic systems if they identify harmful text or tool actions before those actions are generated.
By Max Fomin, Elad David, Amit LeVi
arXiv:2606. 29441v1 Announce Type: cross Abstract: Inference-time safety methods for large language models have proliferated, yet no systematic comparison exists.
By Subhadip Mitra
arXiv:2606. 24952v1 Announce Type: cross Abstract: A central aspiration of mechanistic interpretability is controllability: if we know where a behavior is represented in a model's activations, we should be able to modify it.
By Cosimo Galeone, Anna Ettorre, Minsu Park, Giuseppe Ettorre, Daniele Ligorio
arXiv:2608.02657v2 Announce Type: replace-cross
Abstract: Agentic LLMs are vulnerable to indirect prompt injection (IPI) attacks, e.g., malicious side-tasks hidden in external tool results. While man...
By Jianshuo Dong, Yiming Liu, Maosen Zhang, Nan Deng, Peng Xu, Xiaoping Zhang, Tianwei Zhang, Jie Zhang, Han Qiu