arXiv AI By Liner Xiang, Wenbo Zhang, Hengrui Cai

OTROPE: Optimal Transport-based Robust Off-policy Evaluation for Large Language Models

Read the original on arXiv AI →

The paper introduces OTROPE, a likelihood‑free method for off‑policy evaluation of large language models (LLMs) that uses optimal transport to align labeled samples from a behavior model with unlabeled samples from a target model in a semantic space. OTROPE corrects human‑labeled residuals with proxy predictors, achieving a doubly robust evaluation without requiring behavior‑policy modeling or density‑ratio estimation. The authors provide theoretical guarantees for consistency and convergence, and demonstrate through synthetic and real LLM tasks that OTROPE outperforms existing baselines and can elevate weaker evaluators to match or exceed stronger ones.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
2d ago

Quantifying Behavioral Tails in Black-Box Language Models

The paper introduces RareTrap, a framework that estimates the probability of severe behaviors in black‑box large language models. RareTrap constructs a geometry‑aware mapping from a low‑dimensional latent space into token‑embedding space using a surrogate LLM, creating an explicit and reproducible distribution over input prompts. By applying a response‑level performance function and sequential rare‑event simulation, RareTrap concentrates evaluations on increasingly severe behaviors while preserving probability, enabling estimation of such behaviors with as few as 200 evaluations across multiple open‑weight and frontier models.

By Elsayed Eshra, Ali Al-Lawati, Dongwon Lee, Suhang Wang