arXiv AI

Knowledge-Intensive Video Generation

arXiv:2606. 01285v1 Announce Type: cross Abstract: Text-to-video generation has advanced rapidly in visual quality, but remains under-evaluated for factuality and practical usefulness.

arXiv Computer Vision
Sep 25

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

WanPE is a 397‑B parameter prompt‑enhancement model that learns director‑level cinematic planning from 1.05 M real‑world videos. It generates shot‑level cinematic plans through video‑grounded reverse construction and uses Semantic‑Consistency GRPO (SC‑GRPO) to maintain user intent across shots and time. In evaluations, WanPE improves human preference over raw prompts by up to 50.86 points for 30‑second videos and outperforms commercial offerings for shorter durations.

By Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong
Hugging Face Trending Papers
Aug 10

RAVEN-Eval: Rubric-Guided Automatic Evaluation for AI Video Generation Models Based on LMM Preference Judgement

AI video generation has advanced rapidly and entered widespread commercial use. As a result, quality differences among videos produced by state-of-the-art AI video generation models~(AIVGMs) have become increasingly difficult to discern using conventional evaluation criteria, such as visual fidelity and semantic instruction following.

arXiv Computer Vision
Aug 27

VGA-BenchV2: An Expanded Unified Benchmark and Multi-Model Framework for Evaluating Video Aesthetics and Generation Quality

VGA‑BenchV2 is an expanded, human‑aligned benchmark and optimization framework that jointly evaluates video generation quality and aesthetic value. It builds on the original VGA‑Bench taxonomy, adding 52 sub‑dimensions and 1,016 curated prompts to generate over 60,000 videos from 12 mainstream models. The benchmark significantly enlarges human supervision with 36,000 task‑level annotations and introduces a hybrid evaluator (VAQA‑Net, VTag‑Net, VGQA‑Net) that aligns well with human judgments and can be used as a reward model for reinforcement‑learning fine‑tuning.

By Longteng Jiang, DanDan Zheng, Qianqian Qiao, Heng Huang, Huaye Wang, Yihang Bo, Bao Peng, Jingdong Chen, Jun Zhou, Xin Jin
arXiv Computer Vision
2d ago

VTR-Bench: A Systematic Benchmark for Evaluating Visual Text Rendering in Video Generation

VTR-Bench is a new benchmark designed to evaluate how well video generation models render text within scenes. It includes 300 prompts across five real-world scenarios such as advertisements and scientific videos, and uses an automated pipeline with human alignment to assess text fidelity and scene/motion requirements. Experiments on 11 state‑of‑the‑art models show that even the best performer has a word error rate of 0.250, underscoring widespread challenges in visual text rendering.

By Yu Huang, Jungang Li, Zhiyuan Wang, Yonghua Hei, Song Dai, Jiayu Yang, Deyuan Liu, Xiang Zheng, Xiaoshuang Shi, Hao Cheng, Kaidi Xu
Hugging Face Trending Papers
6d ago

SkillPE: Creativity-Oriented Cinematic Skill Evolution for Text-to-Video Prompt Engineering

SkillPE is a prompt‑engineering framework that evolves reusable cinematic skills from expert‑authored seeds to improve text‑to‑video generation for non‑experts. It encodes shot logic, composition, lighting, sound design, and other filmmaking cues in a fine‑grained format, and uses movie references classified as resonators, dissonants, and divergents to refine skill application and inspire creative alternatives. Experiments on StoryEval and VBench demonstrate up to 1.40‑point gains over the strongest baseline and 0.51 points over seed skills on a 7‑point four‑dimensional evaluation, while remaining competitive on benchmark‑native metrics.

arXiv AI
Sep 21

VidOmni-Bench: A Benchmark for Fine-Grained Video Understanding via Spatio-Temporal Event Verification across Complexity and Duration

VidOmni-Bench is a new benchmark for fine‑grained video understanding that asks models to verify whether each event in dense video captions is supported by the video. It contains 500 videos covering five complexity types and durations from 4 seconds to 90 minutes, and uses human‑verified sentence‑level labels to create hard negatives. Experiments show that Video‑LLMs often hallucinate events, struggle to detect incorrect descriptions, and exhibit varying weaknesses depending on video complexity and duration.

By Changbeen Kim, Junwon Chang, Kipyo Kim, Risa Shinoda, Kuniaki Saito, Donghyun Kim
arXiv Computation and Language
Sep 4

KnowVis: Knowledge-Centric Visual Summarization for Video Lectures

KnowVis is a framework that converts linear video lectures into knowledge‑centric visual narratives. It first extracts a detailed concept map from multimodal video content to identify key and challenging concepts, then builds structured knowledge units and synthesizes engaging visual summaries. The authors also provide a curated dataset of 125 educational videos across 10 disciplines, paired with 1,079 visual summaries, and show through automated evaluations and a human study that KnowVis produces more accurate, clear visuals that reduce cognitive load and improve learning effectiveness and knowledge retention.

By Yi Xu, Yifan Hou, Xiaoyu Zhang