arXiv Computer Vision

Target-Checked Reliability Score Refinement for Video Question Answering

arXiv AI
Sep 11

Scored vs. Generated Readouts in Behavioral Language Models: An Empirical Study of Elicitation Format

The study compares two ways of obtaining predictions from language models fine‑tuned on customer behavior: scoring answer tokens directly versus generating a written rationale and then scoring the resulting answer. Across 13 model‑domain cells covering four retail tasks, scored readouts consistently rank outcomes more accurately than generated readouts, with an AUC improvement ranging from 1.5 to 14.5 points. The authors also find that a third readout—eliciting a probability before any verdict—improves calibration but only when outcome rates are represented in training, and they recommend using generated rationales for interpretability while relying on scored heads for ranking.

By Touchapon Kraisingkorn, Krittin Pachtrachai, Wachiravit Modecrua
arXiv Computer Vision
Aug 25

FIRM-Video: Check Before You Score for Reliable Text-to-Video Reward Modeling

arXiv:2608.21839v1 Announce Type: new Abstract: Reliable reward models are essential for text-to-video evaluation and alignment. However, the trade-off between evaluation accuracy and inference effic...

By Peiyuan Zhang, Xiangyu Zhao, Hongbo Liu, Xiaoxing Hu, Mingxin Liu, Shuran Ma, Yunhang Shen, Jian Hu, Haihan Gao, Haoyu Cao, Xue Yang
Hugging Face Trending Papers
Jun 14

Mitigating Visual Hallucinations in Multimodal Systems through Retrieval-Augmented Reliability-Aware Inference

Multimodal large language models (MLLMs) have demonstrated strong capabilities in vision-language understanding and natural-language response generation. However, these systems can still produce overconfident predictions and hallucination-like outputs, particularly when the visual evidence is weak, ambiguous, or semantically inconsistent.

Hugging Face Trending Papers
Aug 10

RAVEN-Eval: Rubric-Guided Automatic Evaluation for AI Video Generation Models Based on LMM Preference Judgement

AI video generation has advanced rapidly and entered widespread commercial use. As a result, quality differences among videos produced by state-of-the-art AI video generation models~(AIVGMs) have become increasingly difficult to discern using conventional evaluation criteria, such as visual fidelity and semantic instruction following.

arXiv Machine Learning
Jun 2

Perception First: A Frontier Native-Video Model with Self-Consistency for Implicit Video Question Answering

arXiv:2606. 01485v1 Announce Type: cross Abstract: We describe our submission to the VRR Challenge @ CVPR 2026, built on the \emph{ImplicitQA} / \emph{VRR-QA} benchmark~\cite{implicitqa}: multiple-choice video question answering in which answers are deliberately \emph{not} observable in any single frame and must be inferred from spatial layout, motion, depth, viewpoint, causality, and social context across discontinuous frames of creative video.

By Ali Alavi