arXiv Statistics ML By Matthew Francis Dixon

Identification and Honest Recovery from Semantic Observation Kernels: Operator Error, Coarsening, and Stability

Read the original on arXiv Statistics ML →

The paper addresses the challenge of extracting reliable posterior uncertainty from probabilistic text generators, such as large language models, which provide phrase-level probabilities that are prompt-dependent and incomplete. It formulates the recovery of the target posterior as a semiparametric inverse problem and introduces honest recovery guarantees that jointly consider calibration error, measurement noise, incomplete probabilities, and weak identification. Simulations and studies on frozen language models validate the method’s coverage and stability, showing how to determine when a semantic measurement can be trusted for inference or when recalibration or abstention is needed.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Statistics ML.

arXiv AI
Jul 23

Rethinking Uncertainty Evaluation in Large Language Models

arXiv:2607. 19367v1 Announce Type: new Abstract: Calibration is the primary criterion for evaluating LLM confidence, but it is insufficient: it admits trivially incoherent estimators, depends on the evaluation distribution, and does not test the extent to which the estimation can be interpreted as a consistent, underlying probability function.

By Krish Matta, Atharv Naphade, Andy Zou