CalArena: A Large-Scale Post-Hoc Calibration Benchmark
arXiv:2605. 30188v2 Announce Type: replace-cross Abstract: Reliable probability estimates are critical in many machine learning applications, yet modern classifiers are often poorly calibrated.
arXiv:2601. 07965v2 Announce Type: replace Abstract: When a model knows when it does not know, many possibilities emerge.
arXiv:2605. 30188v2 Announce Type: replace-cross Abstract: Reliable probability estimates are critical in many machine learning applications, yet modern classifiers are often poorly calibrated.
The paper introduces Calibration-Aware Uncertainty Cascades (CAUC), a post‑hoc framework that calibrates each model’s confidence independently and uses these calibrated scores to decide when to accept an early prediction, invoke a stronger model, or combine outputs. CAUC establishes a common reliability scale across heterogeneous models, decoupling deployment policies from specific model pools or budgets. Experiments on six language benchmarks show a 1.9% relative accuracy gain over strong‑model‑only inference while cutting strong‑model calls by about 47%, and on image classification it maintains or improves performance while reducing GFLOPs by up to 57%.
arXiv:2603. 20895v3 Announce Type: replace-cross Abstract: Existing routers rely on semantic query features or handcrafted features, which often fail to capture model-specific failures or intrinsic task difficulty.
arXiv:2608.29996v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) achieve strong performance across many multimodal tasks; however, they often exploit spurious object-background...
arXiv:2509.24988v2 Announce Type: replace-cross Abstract: Generating accurate and calibrated confidence estimates is critical for deploying LLMs in high-stakes or user-facing applications, and remain...
arXiv:2605. 23595v2 Announce Type: replace-cross Abstract: The rapid advancement of machine learning has led to an unprecedented expansion of model ecosystems, making it increasingly difficult to assess the reliability of newly released models on unseen and unlabeled data.
ActMap is a new white‑box representation that compresses the entire hidden‑state trajectory of a language model during generation into a fixed 12 × 32 × 128 tensor. This compact 96 KiB map can be captured with no overhead and is read by a lightweight Vision Transformer to estimate answer correctness in a fraction of a millisecond. In experiments on short‑answer QA, math, and summarization, ActMap outperforms sampling, token‑probability, attention, and embedding baselines and matches a larger ACT‑ViT detector while achieving lower calibration error on most test pairs.
The paper argues that calibration—how well a language model’s confidence aligns with its actual correctness—should be a standard evaluation metric for large language models (LLMs). It notes that while calibration metrics exist, they are rarely applied outside specialized NLP subfields, leading to unverified confidence scores in new models, datasets, and benchmarks. The authors highlight the risks of miscalibration both at deployment (overconfident errors causing harm) and in research workflows (affecting LLM-as-a-judge, synthetic data generation, and active learning). They call for every NLP subfield to pair its primary performance metric with a calibration score, treating calibration as an essential property of every model.
The paper investigates how Vision‑Language Models (VLMs) often report high confidence even after self‑correcting or arriving at wrong answers, a phenomenon the authors attribute to the verbalized confidence being largely independent of the model’s reasoning trajectory. By analyzing content variation, token masking, and hesitation markers, the authors demonstrate that confidence does not adequately reflect the actual reasoning process and that calibration training can sometimes worsen this disconnect. To address this blind spot, they introduce the Trajectory‑Grounding Score (TGS) in two forms—TGS‑self and TGS‑pair—and propose TGS‑Bench, a suite of 10 benchmarks that reveal divergences between conventional calibration metrics and trajectory‑grounded confidence.
arXiv:2608. 08700v1 Announce Type: new Abstract: Reliable evaluation of tool routing is critical as Large Language Models increasingly operate as autonomous agents.
The paper introduces DUA-D2C, a Dynamic Uncertainty-Aware Divide2Conquer method that improves overfitting remediation in deep learning. It refines the traditional Divide2Conquer approach by dynamically weighting subset models based on a composite score of accuracy and normalized prediction entropy, allowing the central model to learn more from generalizable and confident edge models. The authors provide theoretical justification, show reduced model variance, and demonstrate significant generalization gains across image, audio, and text benchmarks, even when combined with standard regularizers like Dropout.
arXiv:2606. 07822v1 Announce Type: cross Abstract: As language models improve and become increasingly deployed to solve a variety of tasks, trustworthiness becomes essential.