arXiv Computation and Language

The Dice Roll Method: A Standardized Protocol for Repeated-Query Auditing of Large Language Model Brand Recommendations

The Dice Roll Method is a standardized protocol for auditing large language model brand recommendations through repeated queries. It decomposes total response variance into sampling, prompt‑phrasing, run‑to‑run, and model‑version components, and uses a negative‑binomial mixed model, Cliff’s delta, and bootstrap techniques to guide iteration counts. The study identifies three iteration tiers—exploratory (n=5), confirmatory (n=10), and rigorous (n=15)—and recommends a compact battery of four complementary metrics for robust evaluation.

arXiv AI
Sep 25

How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure

The paper investigates the reliability of ranking tables produced by small-sample evaluations of large language models (LLMs). Using LLM‑inferred prompt structure across eight model variants, the authors find that prompt‑structure recovery is highly unstable, with only the bottom of the ranking consistently reproducible. They demonstrate that standard evaluation practices can misrepresent model performance and propose reporting practices to improve transparency.

By Dipankar Sarkar
arXiv Computation and Language
Sep 25

Measuring Brand and Source Discovery under Repeated LLM Queries: A Finite-Sample Audit

The study audits large language model (LLM) outputs by measuring how well repeated queries recover a collected set of responses versus the full set of possible outputs. Using sample-based rarefaction on 4,500 responses from 50 buying questions across six configurations, the authors find historical-dictionary median recovery rates between 92.6% and 95.2%, which drop to 89.5%–94.7% after re‑adjudicating all candidate strings. Additional analyses with Gemini 3.1 Pro annotations and matched roster data confirm that recovery percentages vary with extraction methods, question selection, and the finite reference collection, underscoring the need for explicit measurement definitions and sensitivity analyses in LLM audits.

By Dmitrij \.Zatuchin
arXiv Machine Learning
Aug 31

The Instability of Safety: How Random Seeds and Temperature Expose Inconsistent LLM Refusal Behavior

The paper challenges the assumption that large language models (LLMs) produce deterministic safety responses by examining how random seeds and temperature settings affect refusal decisions. Across four instruction‑tuned models and 876 harmful prompts, 18‑28% of prompts flipped between refusal and compliance depending on sampling configuration, with higher temperatures reducing decision stability. The authors introduce a Safety Stability Index (SSI) and recommend multi‑sample evaluation protocols that account for stochastic variation rather than relying on single‑shot tests.

By Erik Larsen
arXiv Machine Learning
2d ago

Tail-Influence Sampling for CVaR Policy Evaluation

arXiv:2609.38096v1 Announce Type: new Abstract: Policies with similar mean returns can differ sharply in rare failures, yet estimating lower-tail conditional value-at-risk (CVaR) accurately can requi...

By Pauline Bourigault, Xiaotong Ji, Matthieu Zimmer, Rasul Tutunov, Haitham Bou-Ammar