arXiv:2609.39445v1 Announce Type: cross
Abstract: Time series foundation models (TSFMs) commonly adapt to new data by attaching a single trainable head to a frozen backbone, a one-size-fits-all setup...
By Hung Phan, Thuy T. Nguyen, Minh Ngoc Dinh, Nhat-Quang Tran
arXiv:2608. 13756v1 Announce Type: new Abstract: Two GPU kernels implementing the same scaled INT8 GEMM interface are usually treated as interchangeable.
By Teng-Ruei Chen
arXiv:2608. 11822v1 Announce Type: cross Abstract: A growing body of work reports that language models represent task-relevant latent structure that they fail to use.
By Xining Xun
arXiv:2608. 15787v1 Announce Type: cross Abstract: Two Mixture-of-Experts (MoE) forward passes can share every weight yet route the same token through different experts.
By Cedric Caruzzo, Donggeun Yoo, Tae Soo Kim
arXiv:2608. 03842v1 Announce Type: cross Abstract: When a language model fails on surface-perturbed input (typos, OCR noise, homophones), "which layer is responsible" has three natural operationalizations: where representations diverge most (sensitivity), where restoring clean activations recovers the prediction (causality), and where a small adapter can repair the damage (compensatory capacity) - and we show these three layer maps dissociate.
By Nathan Labiosa, David Buff, Ena Nayak, Erica Donno
arXiv:2606. 10154v1 Announce Type: new Abstract: Quantized checkpoints are often screened first with quality metrics and only later, if at all, with direct safety tests.
By Sahil Kadadekar
arXiv:2609.06473v1 Announce Type: new
Abstract: Inference-time activation steering enables behavioral control of large language models without parameter modification, while post-training quantization...
By Saurav Bhandari, Benjamin Wade
The paper demonstrates that a verifier used in closed‑loop agent debugging can inadvertently reveal the answer it is meant to test, rendering solver comparisons meaningless. In a study of 12 development cases, both exact minimum hitting set and a greedy method returned identical supports, and an audit showed that exact‑anchor predicates always produced the planted fault pair. The authors propose a support‑gated verification contract that requires a clean reference map and runtime evidence before an independently calibrated signal can confirm a detection, and validate this approach on 1,440 held‑out cases with a low false‑admission rate.
By Peiying Zhu, Sidi Chang
arXiv:2606. 10703v1 Announce Type: new Abstract: Interpretability methods routinely use population-level summary statistics over observed model behaviour to license claims about the effects of targeted interventions on specific computations; in Pearl's terms, they treat rung-1 associational evidence as if it supported rung-2 interventional conclusions, a move whose validity is rarely tested.
By Leonard Engmann, Christian Medeiros Adriano, Holger Giese
arXiv:2609.06934v1 Announce Type: cross
Abstract: Post-hoc safety training (RLHF, DPO) is the dominant way to align large language models, yet jailbreaks (Zou et al., 2023b), fine-tuning attacks (Qi...
By Srikanth Malla, Chiho Choi, Joon Hee Choi
arXiv:2609.07901v1 Announce Type: new
Abstract: Weight quantization largely determines the economics of serving open-weight LLMs. Its costs are usually assessed with capability benchmarks, on which 4...
By Dachi Kurtskhalia
arXiv:2607. 29400v1 Announce Type: new Abstract: A routing decision can be revised at the next transaction, but a latched source exclusion persists across later decisions.
By Xiyang Zhang, Hongzhi Wang, Yuanhe Tian