arXiv:2607. 28968v1 Announce Type: cross Abstract: Renault, Bergeaud, and Bosquet (hereafter RBB) argue that dating LLM adoption as the first month in which an author's abstract is flagged induces a stopping-time selection that can produce a positive event-study path even when there is no causal effect.
By Keigo Kusumegi, Xinyu Yang, Paul Ginsparg, Mathijs de Vaan, Toby Stuart, Yian Yin
arXiv:2608. 07914v1 Announce Type: new Abstract: Behavioral contamination detectors can return "no evidence" either because a benchmark is clean or because the audit has little power.
By Ibne Farabi Shihab, Sanjeda Akter, Anuj Sharma
arXiv:2609.07944v1 Announce Type: new
Abstract: Existing causal-inference benchmarks for LLMs mostly score method descriptions or whether generated code runs, not whether the executed workflow recove...
By Yonghong Zhang, Ricardo Correia, Isabel M. Parra, Yong Xie
arXiv:2608.21334v2 Announce Type: replace
Abstract: Short observational pricing panels often contain many observations but few distinct price movements. We evaluate the inferential consequences of th...
By Pedro Cadahia Delgado
arXiv:2606. 01090v1 Announce Type: cross Abstract: Equivariance theory predicts that an architectural symmetry prior reduces sample complexity by a factor of |G|; this is widely cited but rarely measured as a scaling law with controls that separate the prior from its confounds.
By Ahmed M. Adly
arXiv:2503. 20546v2 Announce Type: replace-cross Abstract: We consider the problem of estimating the expected causal effect $E[Y|do(X)]$ for a target variable $Y$ when treatment $X$ is set by intervention, focusing on continuous random variables.
By Marlies Hafer, Alexander Marx
arXiv:2608. 06940v1 Announce Type: new Abstract: LLM judge panels are a standard evaluation tool, but prior work reports highly correlated panel errors: nine judges provide roughly the effective information of two independent ones, and aggregation closes only a small fraction of the gap.
By Yang Shu
arXiv:2605. 09169v2 Announce Type: replace-cross Abstract: A Mamba state-space model trained only for next-step prediction appears to recover Granger-causal structure through a simple readout $S = |W_{out} W_{in}|$, with early experiments suggesting the phenomenon generalized across architectures and benefited from interventional data at $p < 10^{-5}$.
By Ankit Hemant Lade, Sai Krishna Jasti, Indar Kumar, Aman Chadha
Whether a language model has improved itself is increasingly judged not by mean accuracy but by which individual problems it gains and loses. Tracking these transitions means differencing two noisy estimates, leaving them vulnerable to measurement artifacts.
arXiv:2608. 20290v1 Announce Type: new Abstract: Whether a language model has improved itself is increasingly judged not by mean accuracy but by which individual problems it gains and loses.
By Cheng Xu, Nan Yan, Liming Chen, M-Tahar Kechadi
arXiv:2604. 18245v3 Announce Type: replace Abstract: Large language models operate in protocols containing multiple calls, yet added calls are usually evaluated only by their net effect.
By Fernando Reitich
arXiv:2608.21334v1 Announce Type: new
Abstract: Short observational pricing panels can contain many observations while offering only a small number of distinct price movements. This paper studies the...
By Pedro Cadahia Delgado