← Back to all news
arXiv Computation and Language September 23, 2026 By Sophie Henning, Georg Hofmann, Alexander Schulte, Alexander Fraser, Annemarie Friedrich

How to Estimate Whether You Have Found Several Needles in a Haystack: Measuring Calibration in Multi-Label Text Classification

Read the original on arXiv Computation and Language →

The Flow has not summarised this story yet — read it at arXiv Computation and Language.

  • llms
  • nlp
  • benchmarks

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

Hugging Face Trending Papers
Sep 22

How to Estimate Whether You Have Found Several Needles in a Haystack: Measuring Calibration in Multi-Label Text Classification

A key factor in deciding whether to trust an automatic prediction is its confidence score, which should be calibrated to match the actual probability of the prediction being correct. Most confidence c...

llmsnlpbenchmarks
More like this →
arXiv Machine Learning
Jul 15

MLPTR-CC: Multi-label Pathology Test Recommendation using Classifier Chains and SHAP

arXiv:2607. 08299v2 Announce Type: replace Abstract: Diagnostic decision making often relies on a sequence of pathology tests that bridge patient symptoms and final disease diagnosis.

By Abu Rafe Md Jamil, Nayan Malakar
More like this →
arXiv Machine Learning
4d ago

Hierarchical Utility Calibration for Structured Multiclass Decisions

arXiv:2609.36532v1 Announce Type: cross Abstract: In multiclass probabilistic prediction, Utility Calibration (UC), which focuses auditing on specified utilities, has recently received attention as a...

By Futoshi Futami, Jerry Huang, Ichiro Takeuchi
computer-vision
More like this →
arXiv Computer Vision
3d ago

SCALE: Synthetic Calibration via Agreement Labeling in Embedding Space

arXiv:2609.38705v1 Announce Type: new Abstract: Foundation models for computational pathology are usually evaluated using AUC and accuracy, while calibration is often left untested. This matters beca...

By Wenjun Liu, Saeed Hassanpour
rag
More like this →
arXiv Machine Learning
Sep 22

SupportCal: Label-Free Calibration of Post-Trained LLMs via Reference Support and Corroboration

arXiv:2609.24303v1 Announce Type: new Abstract: Post-training often improves task performance but can degrade confidence calibration, leaving post-trained language models (PoLMs) more overconfident t...

By Linhan Luo, Lequan Lin, Dai Shi, Feng Chen, Jos\'e Miguel Hern\'andez-Lobato, Junbin Gao
llmssafety
More like this →
arXiv Machine Learning
Jul 16

Temperature Scaling Is Not Enough: Calibration Gaps Under Human Label Distributions

arXiv:2607. 13423v1 Announce Type: new Abstract: Temperature scaling is the dominant post-hoc calibration method in modern deep learning.

By Wisdom Dogah
safety
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea