arXiv Computation and Language By Daniel Yoo, Adrians Skapars

Probe Generalization as Subspace Selection for OOD Deception Detection

Read the original on arXiv Computation and Language →

Linear probes can identify behaviors in language model activations but often fail on out‑of‑distribution data. This study shows that projecting inputs onto a small set of principal components (PCs) from the training distribution allows probes for Llama‑3.1‑8B‑Instruct to transfer across three deception‑detection datasets, nearly matching probes trained directly on the test data. By scoring PCs with an LLM judge to select those that encode transferable deception directions, the authors close the baseline‑to‑oracle gap by 78% on Insider Trading Report and 25% on Sandbagging, revealing that subspace selection largely determines OOD robustness.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.