Hugging Face Trending Papers

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification

Confidence estimation is essential when LLMs are used for classification, indicating when predictions can be trusted. However, common approaches such as verbalization produce extremely sparse outputs.

arXiv AI
Aug 11

Embedding Trust: Semantic Isotropy Predicts Nonfactuality in Long-Form Text Generation

arXiv:2510. 21891v2 Announce Type: replace-cross Abstract: To deploy large language models (LLMs) in high-stakes application domains that require substantively accurate responses to open-ended prompts, we need reliable, computationally inexpensive methods that assess the trustworthiness of long-form responses generated by LLMs.

By Dhrupad Bhardwaj, Julia Kempe, Tim G. J. Rudner