Belief-reality separation lives in routing over a shared value slot in language models
Read the original on arXiv Machine Learning →arXiv:2607. 11945v1 Announce Type: cross Abstract: Capable language models hold what a character believes apart from what is true: told "Anna believes the cup is blue; in reality it is red," they answer blue about Anna and red about the world.
Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.