arXiv AI By Maheep Chaudhary, Fazl Barez

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors

Read the original on arXiv AI →

arXiv:2505. 14300v2 Announce Type: replace Abstract: White-box monitoring is increasingly adopted as an auditing tool as Large Language Models (LLMs) are deployed in daily operations to ensure safe model behavior.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 9

When Behavioral Safety Evaluation Fails: A Representation-Level Perspective

arXiv:2606. 08044v1 Announce Type: cross Abstract: Large Language Model (LLM) safety has often been evaluated at the behavior level, which provides limited evidence of internal robustness, as these evaluations target outputs rather than representation-level vulnerability under intervention.

By Enyi Jiang, Anders Gj{\o}lbye, Yibo Jacky Zhang, Sanmi Koyejo
arXiv Machine Learning
Sep 3

The Implications of Linguistic Illegibility for LLM Security

The paper introduces the concept of "linguistic illegibility," describing how a large language model’s (LLM) language outputs and extracted linguistic features may not accurately reflect its internal computations. It argues that because LLMs compute primarily in activation spaces, any reliance on linguistic self‑reporting for security—such as chain‑of‑thought monitoring or constitutional self‑critique—cannot be fully reliable. The authors propose taint tracking and other sandboxing techniques that do not depend on the model’s linguistic state as a more robust security foundation.

By James Mickens