arXiv AI

Hidden APIs in Language Models: Discovering Reusable Causal Interfaces from Forked Futures

arXiv AI
Sep 12

Portable Semantics, Private Dialects: Reuse and Negative Transfer in Latent Communication Between Language-Model Cells

The study investigates whether independently trained language‑model societies share a common packet language and how inherited interface states affect learning. A comprehensive audit of 30 pairwise interactions among six restricted societies shows that only one pair is fully interoperable, another is partially compatible, and the remaining 26 pairs fail across all alignment levels. Further experiments reveal that a globally trained communication interface can act as a severe negative‑transfer prior, but inherited interfaces never outperform fresh‑interface controls by the preregistered margin.

By Narcis Marincat
Hugging Face Trending Papers
Sep 10

Portable Semantics, Private Dialects: Reuse and Negative Transfer in Latent Communication Between Language-Model Cells

The study investigates whether independently trained language‑model societies share a common packet language and how inherited interface states affect learning. A comprehensive audit of 30 pairwise interactions among six restricted societies shows that most cross‑initialization pairs fail to interoperate, with only one pair achieving full bidirectional compatibility. Further experiments reveal that reinitializing only the packet reader, writer, and mouth dramatically improves accuracy, while inherited interfaces never outperform fresh ones by the preregistered margin.

arXiv Machine Learning
Jun 5

Pattern Selectivity is Not Task-Causal Structure: A Cross-Architecture Mechanistic Study of Composed-Task Circuits in 1B-Class Language Models

arXiv:2606. 05378v1 Announce Type: new Abstract: We test whether a single screen-and-ablate recipe -- identify attention-head circuits by task-pattern selectivity, then verify by causal ablation against a matched-random null -- produces consistent mechanistic claims across model families.

By Yongzhong Xu
arXiv Machine Learning
Aug 5

Sensitivity, Causality, and Repair Dissociate: A Layer-Wise Analysis of Perturbation Robustness and Its Scaling

arXiv:2608. 03842v1 Announce Type: cross Abstract: When a language model fails on surface-perturbed input (typos, OCR noise, homophones), "which layer is responsible" has three natural operationalizations: where representations diverge most (sensitivity), where restoring clean activations recovers the prediction (causality), and where a small adapter can repair the damage (compensatory capacity) - and we show these three layer maps dissociate.

By Nathan Labiosa, David Buff, Ena Nayak, Erica Donno
arXiv Computation and Language
Aug 28

SCIT: Testing Causal Cache Carriers in Latent Chain-of-Thought Models

SCIT (Suffix Cache Interchange Test) is a causal protocol designed to identify which transformer components carry counterfactual computations in latent chain-of-thought models. By constructing exact source‑recipient counterfactuals and applying sufficiency tests, K/V splits, hidden‑state controls, and semantic source controls, SCIT demonstrates that counterfactual arithmetic primarily transfers through value‑cache suffix trajectories rather than hidden states or keys. The method reveals carrier‑regime shifts across different GPT‑2 checkpoints, providing a cache‑level diagnostic and a competence‑gated carrier map for arithmetic mechanisms.

By Yi Ding, Lijun Huang, Menglin Yang