The Truth Was Never Gone: Perfect Aliasing in Compliant-Context Truth Probes
Read the original on arXiv Machine Learning →The paper introduces the concept of perfect aliasing, where a truth probe that aligns truthful reporting with a task’s prescribed action cannot differentiate between the two based solely on its labels. In a binary reporting game, probes fitted on compliant contexts yield identical optimizations, while on rival contexts their labels are complementary, causing their AUROCs to sum to one across 751 cell-layer pairs. By employing randomized codebooks and mixed-context fitting, the authors demonstrate that separating prescribed output symbols from semantic action enables perfect recovery of truth, achieving an AUROC of 1.000 on rival trials for a reward-trained Gemma-2-9B policy, whereas conventional probes perform near chance.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.