arXiv Machine Learning By Oliver Steele, Jiangtao Wen, Yuxing Han

Belief-reality separation lives in routing over a shared value slot in language models

Read the original on arXiv Machine Learning →

arXiv:2607. 11945v1 Announce Type: cross Abstract: Capable language models hold what a character believes apart from what is true: told "Anna believes the cup is blue; in reality it is red," they answer blue about Anna and red about the world.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jul 20

Verbalizable Representations Form a Global Workspace in Language Models

arXiv:2607. 15495v1 Announce Type: cross Abstract: Out of everything the human brain processes, only a small fraction is consciously accessible, in the sense of being available for verbal report, deliberate control, and flexible reasoning.

By Wes Gurnee, Nicholas Sofroniew, Adam Pearce, Mateusz Piotrowski, Isaac Kauvar, Runjin Chen, Anna Soligo, Paul Bogdan, Euan Ong, Rowan Wang, Ben Thompson, David Abrahams, Subhash Kantamneni, Emmanuel Ameisen, Joshua Batson, Jack Lindsey
arXiv AI
Aug 20

Intercepting the Kangaroo: Experimental Astrolinguistics with Constructed Lexicons, Active Probing, and Large Language Models as Informants and Hypothesis Proposers

The study turns the speculative field of astrolinguistics into an experiment by using two large language models with deliberately incompatible constructed lexicons as informants. A scripted orchestrator translates between the two category systems, and a protocol combining cross‑situational elimination, predictive probes, active scene selection, and a stricter recovery round successfully prevents the ‘kangaroo effect’—the silent attachment of a word to the wrong referent—in over 400 simulated and live runs. When informant noise is introduced, the protocol remains robust up to 2% per‑word noise and largely abstains rather than errs at higher noise levels, while a generate‑and‑test loop allows recovery of words outside the scripted hypothesis space, achieving full coverage as the rule‑proposing LLM’s capability increases. whyItMatters":"The protocol demonstrates that experimental astrolinguistics can reliably avoid mistranslations and recover unknown terms, showing that correctness is governed by the protocol while coverage depends on the instruments used."

By Francesco Cordella, Mauro Cappelli