arXiv AI By Rico Angell, Raghav Singhal, Zachary Horvitz, Zhou Yu, Rajesh Ranganath, Kathleen McKeown, He He

Estimating Tail Risks in Language Model Output Distributions

Read the original on arXiv AI →

arXiv:2604. 22167v2 Announce Type: replace-cross Abstract: Language models are increasingly capable and are being rapidly deployed on a population-level scale.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
4d ago

Quantifying Behavioral Tails in Black-Box Language Models

The paper introduces RareTrap, a framework that estimates the probability of severe behaviors in black‑box large language models. RareTrap constructs a geometry‑aware mapping from a low‑dimensional latent space into token‑embedding space using a surrogate LLM, creating an explicit and reproducible distribution over input prompts. By applying a response‑level performance function and sequential rare‑event simulation, RareTrap concentrates evaluations on increasingly severe behaviors while preserving probability, enabling estimation of such behaviors with as few as 200 evaluations across multiple open‑weight and frontier models.

By Elsayed Eshra, Ali Al-Lawati, Dongwon Lee, Suhang Wang