arXiv AI By Andrew Mack, Nina Panickssery, Alexander Matt Turner

Mechanistically Eliciting Latent Behaviors in Language Models

Read the original on arXiv AI →

arXiv:2606. 29604v1 Announce Type: cross Abstract: We aim to discover diverse, generalizable perturbations of LLM internals that can surface hidden behavioral modes.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.