arXiv AI

ReliableTableQA:How Much Supervision Does Reliability Annotation Need?

arXiv:2607. 20537v1 Announce Type: cross Abstract: We introduce ReliableTableQA, a framework for training an LLM to annotate the statistical reliability of tabular QA results, not whether the query is answerable, but whether the computed answer is statistically meaningful.

arXiv Machine Learning
Aug 21

LODESTAR: Robust Entropy-Based Answer Selection in Retrieval-Augmented Generation for Question Answering -- Directing Frozen-LLM Entropy with a Reinforcement-Learned Prompt Polarizer under Misleading Passages

arXiv:2608. 11922v2 Announce Type: replace-cross Abstract: Predictive-distribution entropy is a strong answer-selection rule in retrieval-augmented generation (RAG) for question answering: across five QA benchmarks, selecting the answer a frozen respondent LLM produces with the lowest answer-token entropy lifts mean $F_1$ from 0.

By Hung-Chun Hsu, Po-Jen Ko, Che-Cheng Wu, Li-Yang Chang, Chuan-Ju Wang
Hugging Face Trending Papers
6d ago

RGDT-Bench: Benchmarking LLM Reasoning for Rule-Governed Decisions and Their Justifications

RGDT-Bench is a new benchmark that evaluates large language models on Rule‑Governed Decision Tasks, where models must apply external rules to facts, justify decisions, and provide checkable justifications. The benchmark offers 202.1K condition‑level supervision slots across four task tracks and eight task‑probe combinations, and it labels warrant completeness through label‑blind extraction and deterministic checks. Evaluation shows that among correct responses, 40.2% of warrants are incomplete, and existing evaluators struggle to detect this, prompting the authors to train a reward model that improves AUROC to 69.24% and outperforms outcome‑supervised baselines.

arXiv Machine Learning
Sep 17

Safety-Flag: A Unified Benchmark for the Reliability and Calibration of LLM Content Moderators

Safety-Flag is a unified benchmark that consolidates seven popular safety datasets into a single balanced flag/do‑not‑flag protocol, providing item‑level decisions and confidence scores for multiple large language models and dedicated guards. The benchmark evaluates moderator reliability across three dimensions—error direction, probability calibration, and confidence‑based error ranking—revealing that aggregate accuracy masks significant differences, such as one model flagging 85% of benign content while another misses 54% of harmful content. The study shows that general‑purpose models are overconfident, but temperature tuning can substantially improve calibration, and confidence‑based abstention can reduce selective risk, though performance varies with how well confidence ranks errors.

By Yibo Hu