arXiv Machine Learning

Quantifying Behavioral Tails in Black-Box Language Models

The paper introduces RareTrap, a framework that estimates the probability of severe behaviors in black‑box large language models. RareTrap constructs a geometry‑aware mapping from a low‑dimensional latent space into token‑embedding space using a surrogate LLM, creating an explicit and reproducible distribution over input prompts. By applying a response‑level performance function and sequential rare‑event simulation, RareTrap concentrates evaluations on increasingly severe behaviors while preserving probability, enabling estimation of such behaviors with as few as 200 evaluations across multiple open‑weight and frontier models.

arXiv AI
2d ago

OTROPE: Optimal Transport-based Robust Off-policy Evaluation for Large Language Models

The paper introduces OTROPE, a likelihood‑free method for off‑policy evaluation of large language models (LLMs) that uses optimal transport to align labeled samples from a behavior model with unlabeled samples from a target model in a semantic space. OTROPE corrects human‑labeled residuals with proxy predictors, achieving a doubly robust evaluation without requiring behavior‑policy modeling or density‑ratio estimation. The authors provide theoretical guarantees for consistency and convergence, and demonstrate through synthetic and real LLM tasks that OTROPE outperforms existing baselines and can elevate weaker evaluators to match or exceed stronger ones.

By Liner Xiang, Wenbo Zhang, Hengrui Cai
Hugging Face Trending Papers
Jul 2

Safety Targeted Embedding Exploit via Refinement

Safety training for large language models (LLMs) is conducted predominantly in English, leaving uncertain how well safety mechanisms generalize to low-resource languages and mixed-language code-switching. We show that this creates an epistemic gap in which models confidently generate harmful responses for inputs that fall outside the distribution of their safety training.

arXiv AI
Sep 23

Pinocchio: Fast Uncertainty Estimates for Black-Box Language Models

Pinocchio is an external calibrator that provides fast uncertainty estimates for black-box language models. It predicts the correctness of responses from seven trained LLMs with an AUROC of 0.862 and can transfer zero‑shot to thirteen unseen models from eight organizations. The method requires only a single forward pass and no access to the target model’s internal states, and a lightweight 0.8B checkpoint achieves comparable performance.

By Kevin David Hayes, Arka Pal, Haosong Zhang, Tom Goldstein, Micah Goldblum
arXiv Machine Learning
Aug 4

The Illusion of Stochasticity in LLMs

arXiv:2604. 06543v2 Announce Type: replace-cross Abstract: In this work, we demonstrate that reliable stochastic sampling is a fundamental yet unfulfilled requirement for Large Language Models (LLMs) operating as agents.

By Xiangming Gu, Soham De, Michalis Titsias, Larisa Markeeva, Petar Veli\v{c}kovi\'c, Razvan Pascanu
arXiv Machine Learning
Sep 22

Logits are All We Need to Adapt Closed Models

arXiv:2502.06806v5 Announce Type: replace Abstract: Many commercial Large Language Models (LLMs) are often closed-source, limiting developers to prompt tuning for aligning content generation with spe...

By Gaurush Hiranandani, Haolun Wu, Subhojyoti Mukherjee, Sanmi Koyejo