arXiv Machine Learning

Quality Is Not a Safety Proxy Under Quantization

arXiv:2606. 10154v1 Announce Type: new Abstract: Quantized checkpoints are often screened first with quality metrics and only later, if at all, with direct safety tests.

arXiv Machine Learning
Sep 22

The Effect of Quantization on Clinical Benchmarks: Accuracy and Safety Across Model Families

The study evaluates how quantization affects accuracy and safety of five 7‑8B language models on clinical benchmarks. INT8 GPTQ shows minimal degradation (≤1.9%) across tasks, while INT4 causes substantial, model‑dependent drops, especially in high‑risk scenarios and safety metrics. Recovery methods such as clinical calibration substitution and QLoRA fine‑tuning yield mixed results, underscoring the need for task‑specific validation.

By Leonard Twagirayezu, Prasenjit Mitra
arXiv AI
Sep 1

Decomposing Wrong-Consensus Agreement in LLM Self-Consistency

The paper investigates the nature of agreement among repeated samples of large language models (LLMs), showing that strong agreement can arise even for incorrect answers. It introduces a pluralistic agreement index, Gamma, which is decomposed into a mechanical component driven solely by per‑case answer preferences and a residual component that captures preference‑unexplained agreement. Experiments on GPT‑4.1 and several open‑weight models demonstrate that mechanical agreement dominates in many settings, while the residual varies with benchmark type and sampling protocol.

By Lizhuo Zhang, Mengmeng Tang, Chenfeng Long, Xiaoyong Tang, Xiang Luo
arXiv Machine Learning
Aug 31

The Instability of Safety: How Random Seeds and Temperature Expose Inconsistent LLM Refusal Behavior

The paper challenges the assumption that large language models (LLMs) produce deterministic safety responses by examining how random seeds and temperature settings affect refusal decisions. Across four instruction‑tuned models and 876 harmful prompts, 18‑28% of prompts flipped between refusal and compliance depending on sampling configuration, with higher temperatures reducing decision stability. The authors introduce a Safety Stability Index (SSI) and recommend multi‑sample evaluation protocols that account for stochastic variation rather than relying on single‑shot tests.

By Erik Larsen
arXiv Machine Learning
Sep 17

No Usable Linear "Capitulation Direction" in Two Small LLMs: A Validation Protocol for Activation-Steering Claims, and a Cross-Family Behavioral Study of Sycophancy Under Pushback

The study examines how two small instruction‑tuned language models, Qwen2.5‑1.5B and Llama‑3.2‑1B, respond to user pushback on TriviaQA. When initially correct, the models flip to a wrong answer in about 42–43% of cases, with the effectiveness of different pushback styles varying by model. Attempts to decode capitulation from the pre‑response residual stream fail under a rigorous validation protocol, revealing overfitting and a measurement hazard that underestimates capitulation by 18–24 percentage points.

By Saad Aamir, Muhammad Awais Bin Adil