Target-Checked Reliability Score Refinement for Video Question Answering
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2609.09184v1 Announce Type: new Abstract: Vision-language model (VLM) confidence may change in aggregate when visual evidence is degraded while remaining structurally inconsistent within indivi...
arXiv:2609.14284v1 Announce Type: new Abstract: Criterion-level grading connects examination performance to learning outcomes, but manual marking introduces workload and variation between markers. Th...
arXiv:2608. 09111v1 Announce Type: new Abstract: AI video generation has advanced rapidly and entered widespread commercial use.
arXiv:2608. 10406v1 Announce Type: cross Abstract: Web search, product search, and question-answering retrieval systems often assign a relevance label and confidence score to each query-candidate pair.
The study compares two ways of obtaining predictions from language models fine‑tuned on customer behavior: scoring answer tokens directly versus generating a written rationale and then scoring the resulting answer. Across 13 model‑domain cells covering four retail tasks, scored readouts consistently rank outcomes more accurately than generated readouts, with an AUC improvement ranging from 1.5 to 14.5 points. The authors also find that a third readout—eliciting a probability before any verdict—improves calibration but only when outcome rates are represented in training, and they recommend using generated rationales for interpretability while relying on scored heads for ranking.
arXiv:2608.21839v1 Announce Type: new Abstract: Reliable reward models are essential for text-to-video evaluation and alignment. However, the trade-off between evaluation accuracy and inference effic...