arXiv AI By SiYuan Ma, Yiqin Luo, Zhangji, Canran Xiao, Albert Gao, Wei Wang, Qiwei Wu, Xinran Li, Jinfeng Wei, Qixin Zhang

Hidden APIs in Language Models: Discovering Reusable Causal Interfaces from Forked Futures

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv AI
Sep 12

Portable Semantics, Private Dialects: Reuse and Negative Transfer in Latent Communication Between Language-Model Cells

The study investigates whether independently trained language‑model societies share a common packet language and how inherited interface states affect learning. A comprehensive audit of 30 pairwise interactions among six restricted societies shows that only one pair is fully interoperable, another is partially compatible, and the remaining 26 pairs fail across all alignment levels. Further experiments reveal that a globally trained communication interface can act as a severe negative‑transfer prior, but inherited interfaces never outperform fresh‑interface controls by the preregistered margin.

By Narcis Marincat
Hugging Face Trending Papers
Sep 10

Portable Semantics, Private Dialects: Reuse and Negative Transfer in Latent Communication Between Language-Model Cells

The study investigates whether independently trained language‑model societies share a common packet language and how inherited interface states affect learning. A comprehensive audit of 30 pairwise interactions among six restricted societies shows that most cross‑initialization pairs fail to interoperate, with only one pair achieving full bidirectional compatibility. Further experiments reveal that reinitializing only the packet reader, writer, and mouth dramatically improves accuracy, while inherited interfaces never outperform fresh ones by the preregistered margin.

arXiv Machine Learning
Jun 5

Pattern Selectivity is Not Task-Causal Structure: A Cross-Architecture Mechanistic Study of Composed-Task Circuits in 1B-Class Language Models

arXiv:2606. 05378v1 Announce Type: new Abstract: We test whether a single screen-and-ablate recipe -- identify attention-head circuits by task-pattern selectivity, then verify by causal ablation against a matched-random null -- produces consistent mechanistic claims across model families.

By Yongzhong Xu