arXiv:2606. 10703v1 Announce Type: new Abstract: Interpretability methods routinely use population-level summary statistics over observed model behaviour to license claims about the effects of targeted interventions on specific computations; in Pearl's terms, they treat rung-1 associational evidence as if it supported rung-2 interventional conclusions, a move whose validity is rarely tested.
By Leonard Engmann, Christian Medeiros Adriano, Holger Giese
arXiv:2606. 27383v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as research assistants, yet it remains unclear whether they can calibrate research takeaways to the strength and scope of the supporting evidence.
By Yu Fu, Yongqi Kang, Yong Zhao
The paper investigates when auxiliary context can genuinely improve multi‑modal time series forecasting. It identifies two necessary dataset‑level conditions: the target must not be dominated by a last‑value shortcut (low autocorrelation) and the context must provide additional information beyond history (non‑zero conditional mutual information). Experiments on a large mixture‑of‑experts model and several fusion mechanisms show that only when both conditions hold does context routing yield a substantial reduction in mean‑squared error; otherwise its contribution collapses to a capacity floor.
By Ruizhe Zhou, Gaoyuan Du, Xiaoyang Liu, Haoqi Yao, Deepayan Chakrabarti, Jiating Lin, Yixuan Shen
arXiv:2606. 03780v1 Announce Type: cross Abstract: Causal tracing of factual recall has been studied predominantly in dense transformer language models, where interventions localize information flow to layers or feed-forward modules.
By Yuetian Lu, Ali Modarressi, Yihong Liu, Hinrich Sch\"utze
arXiv:2607. 24562v1 Announce Type: new Abstract: Large language models serve heterogeneous populations structured by domain, topic difficulty, and linguistic style.
By Murilo Salem, Lu\'isa B\"ohm, Daniel Pontes, Anderson Ferrugem
arXiv:2606. 05378v1 Announce Type: new Abstract: We test whether a single screen-and-ablate recipe -- identify attention-head circuits by task-pattern selectivity, then verify by causal ablation against a matched-random null -- produces consistent mechanistic claims across model families.
By Yongzhong Xu
arXiv:2608. 07183v1 Announce Type: new Abstract: Multimodal fusion architectures typically assume all modalities are available at inference, yet sensor failures, acquisition variability, and cost constraints routinely produce incomplete observations.
By Alireza Moayedikia
The paper demonstrates that a single-direction white‑box attack, which projects a ‘refusal direction’ from a language model’s weights, remains effective against a 320B‑parameter mixture‑of‑experts (MoE) model (GLM‑5.3‑Flash). The attack requires only a few hundred contrastive prompts and no gradient training, and it reduces refusal behavior by up to 89 percentage points across seven harmful‑content benchmarks while leaving overall capability unchanged. However, the attack’s impact is distributed across multiple components—attention, dense, and routed‑expert writers—so that only a joint intervention removes most of the refusal ability, and the conventional module‑name matching approach fails to capture this effect in MoE architectures.
By Yi Shi, Tanyu Chen, Kai Shen
arXiv:2607. 28308v1 Announce Type: new Abstract: Sparse mixture-of-experts (MoE) language models route each token to multiple experts, suggesting a geometric account of their benefit: co-selected experts should contribute distinct representation directions.
By Huiyuan Tian, Bonan Xu, Shijian Li
The paper investigates whether confidence signals from fine‑tuned large language models can improve extractive question answering that relies heavily on retrieval. Experiments on four 7‑9B model families show that retrieval alone recovers 92–99.8% of the best possible accuracy, leaving little room for confidence‑based routing or adaptation to help. The sequence‑likelihood confidence metric, even after recalibration or temperature scaling, fails to provide a statistically significant benefit across different correctness criteria and answer lengths, and the study ultimately offers a set of pre‑specified negatives with explicit dependencies as its main contribution.
By Gunwoo Lee, Changmin Sung, Sang-Hwan Gwak, Ina Kim, Ji-Young Choi, Kyong-Ha Lee
arXiv:2606. 19625v2 Announce Type: replace-cross Abstract: We use training-data attribution as an interpretable tool for capability discovery, mapping which regions of the pretraining corpus support social-reasoning versus STEM-reasoning in OLMo3-7B.
By Glenn Matlin, Chandreyi Chakraborty, Saehee Eom, Mika Okamoto, Rayan Castilla, Louis Jaburi, Alvin Deng, Taywon Min, Lucia Quirke, Stella Biderman, Mark Riedl
Sparse mixture-of-experts (MoE) language models route each token to multiple experts, suggesting a geometric account of their benefit: co-selected experts should contribute distinct representation directions. Existing evidence often conflates route coherence, candidate quality, and candidate-by-context interaction.