arXiv AI By Youngjae Cho, Won Young Jhoo, Jongsuk Kim

Beyond Scalar IoU: Structured Verification from Rollout Groups for Video Temporal Grounding

Read the original on arXiv AI →

The paper introduces SUTURE, a structured verification method for video temporal grounding that leverages the joint structure of rollout groups rather than scoring each rollout independently. SUTURE conditions verification on the entire rollout group, using disagreement across rollouts to reweight targets and coverage at each position to redistribute reward mass. Experiments on five benchmarks show that SUTURE consistently improves grounding performance across all IoU thresholds and reduces video-start anchoring in reasoning traces, indicating that the joint structure of a rollout group can provide a more informative temporal verifier.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 21

VISD: Enhancing Video Reasoning via Structured Self-Distillation

arXiv:2605. 06094v5 Announce Type: replace-cross Abstract: Training VideoLLMs for complex reasoning remains challenging due to sparse sequence level rewards and the lack of fine grained credit assignment over long, temporally grounded reasoning trajectories.

By Hao Lin, Kunyang Lv, Xu Jiang, Jingqi Tian, Zhongjing Du, Jiayu Ding, Qiaoman Zhang, Hongbo Jin
arXiv AI
Sep 28

CaC: Advancing Video Reward Models via Hierarchical Spatiotemporal Concentrating

The paper introduces CaC, a coarse‑to‑fine anomaly reward model that uses Vision‑Language Models to first scan globally for anomalous time windows, then ground anomalies spatially, and finally reason with structured spatiotemporal Chain‑of‑Thought. It builds the first large‑scale generated video anomaly dataset with detailed annotations and trains the model through a three‑stage progressive paradigm, including reinforcement learning with Group Relative Policy Optimization. Experiments show CaC improves fine‑grained anomaly detection by 25.7% and reduces generated‑video anomalies by 11.7% while enhancing overall video quality.

By Jiyuan Wang, Huan Ouyang, Jiuzhou Lin, Chunyu Lin, Dewen Fan, Boheng Zhang, Haonan Fan, Honglie Wang, Yiyang Fan, Zhenlong Yuan, Zijun Li, Yongrui Heng, Guosheng Lin, Fan Yang