Test-time Calibration Learning for Large Language Model Reasoning
Read the original on arXiv AI →The paper introduces Test-Time Calibration Learning (TTCL), a label‑free framework that adapts both reasoning accuracy and verbalized confidence of large language models directly on unlabeled target‑task data. TTCL generates self‑supervision signals from multiple model responses, enabling calibration without ground‑truth labels and proving theoretically as a bounded surrogate for the ideal calibration objective. Experiments on mathematical reasoning and factual question answering show consistent improvements, with base models gaining an average 40.13% in accuracy and 70.80% in ECE reduction across eight benchmarks, and further gains under domain shift.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.