arXiv:2609. 31181v1 Announce Type: new Abstract: Black-box model identification works by scoring a model's response to natural-language prompts.
By Nicol\'as Vera Z\'u\~niga
The paper investigates how recursive contamination—retraining language models on their own generated text—affects output diversity across 13 publicly released checkpoints. Using a fixed contamination protocol over five generations, the authors find a wide spread in 4‑gram diversity (0.187 to 0.940), indicating that some models collapse into repetitive fragments while others remain largely unaffected. The study shows that a model’s susceptibility to collapse is an intrinsic property of the checkpoint, not predicted by parameter scale or static indicators, and that simple interventions such as tightening top‑p sampling can significantly slow or halt collapse.
By Yangze Liu, Zhongyi Han
The paper evaluates the effectiveness of tolerance‑based conformance tests for INT8 quantized GEMM kernels used in large language models. By injecting nine faults into a Qwen3‑1.7B reference pipeline, the authors show that most faults shift outputs by at most one bfloat16 spacing, rendering a tolerance of one spacing blind to these errors. They further demonstrate that requantizing weight scales to the nearest power of two aligns CUTLASS and Triton implementations bit‑for‑bit and produces identical token sequences, with only minor perplexity changes.
By Teng-Ruei Chen
The paper introduces AgentDiff, a metric that quantifies how much LLM agents’ answers differ when inputs are altered by meaning‑bearing rewrites (paraphrases, synonym substitutions) versus presentation changes (reordering, formatting, distractors). Across 68 model–benchmark–scaffold combinations involving ten LLMs and over 1,500 questions, meaning‑bearing rewrites consistently produce a roughly 20‑percentage‑point higher inconsistency rate than presentation changes, a gap that persists across severity proxies and remains significant even outside the Qwen family. Trace analysis reveals that meaning‑bearing rewrites preserve the first action but reduce thought similarity from the second step onward, extending the divergence cascade—a phenomenon termed “stealth divergence.”
By Liyun Zhang, Jiayi Guo
arXiv:2511.00763v3 Announce Type: replace
Abstract: We investigate the performance of large language models (LLMs) on repetitive deterministic prediction tasks and study how the sequence accuracy rat...
By Wanda Hou, Leon Zhou, Hong-Ye Hu, Yubei Chen, Yi-Zhuang You, Xiao-Liang Qi
arXiv:2609.14754v1 Announce Type: cross
Abstract: Causal claims about large language model (LLM) internals rest on measurements. Those might include a projection, a cosine, an ablation delta, or an i...
By Orion Reblitz-Richardson