arXiv Computation and Language By Yi Ding, Lijun Huang, Menglin Yang

SCIT: Testing Causal Cache Carriers in Latent Chain-of-Thought Models

Read the original on arXiv Computation and Language →

SCIT (Suffix Cache Interchange Test) is a causal protocol designed to identify which transformer components carry counterfactual computations in latent chain-of-thought models. By constructing exact source‑recipient counterfactuals and applying sufficiency tests, K/V splits, hidden‑state controls, and semantic source controls, SCIT demonstrates that counterfactual arithmetic primarily transfers through value‑cache suffix trajectories rather than hidden states or keys. The method reveals carrier‑regime shifts across different GPT‑2 checkpoints, providing a cache‑level diagnostic and a competence‑gated carrier map for arithmetic mechanisms.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Aug 24

Temporal Validity on Real Software Histories: Eliminating Stale-Fact Errors in Code-Assistant Memory over GitHub Fixes

The paper evaluates a deterministic supersession memory, MemStrata, for retrieval‑augmented generation (RAG) systems on real software history. Using 707 GitHub issues, the authors extracted 130 clean atomic state transitions where a single value changes from pre‑fix to post‑fix. MemStrata achieved 0.91 answer accuracy versus 0.57–0.59 for standard RAG, eliminating stale‑fact errors that RAG returned 36–38% of the time, while maintaining comparable retrieval latency.

By Neeraj Yadav
arXiv AI
Sep 24

Are Stated Reasoning Steps Causally Load-Bearing?

The study investigates whether the reasoning steps a language model writes are causally responsible for its answers. Using a causal intervention method on the activation stream, the authors find that for Qwen3-4B, about 77% of stated steps are causally load‑bearing, while behavioral tests overestimate this by roughly 11 percentage points. The faithfulness of reasoning decreases with model size and depth of reasoning, especially for the smaller Qwen3-1.7B.

By Abhiram Bhupatiraju, Rayan Nyaupane