arXiv Machine Learning By Luca Zhou

Tail-Shape Estimation in LLM Evaluation Is Fragile: A Protocol for Diagnosing False Positives

Read the original on arXiv Machine Learning →

arXiv:2606. 16511v1 Announce Type: new Abstract: Recent work motivates moving large language model (LLM) evaluation from mean-based to tail-aware metrics, including conditional value-at-risk and tail-index estimates of reward-model error.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.