arXiv:2609.24083v1 Announce Type: new
Abstract: Generative AI enables scalable production of educational videos, but current systems largely focus on producing visually coherent content rather than s...
By Xinchen Ma, Shuimu Wang, Gaole He, Yanbin Zhang, Chunyang Wang, Yunshi Lan, Weining Qian
arXiv:2608. 08852v1 Announce Type: new Abstract: AI agents can now solve problems, answer like subject experts, and generate long-form multimodal content.
By Yi-Cheng Lin, Yu-Kai Guo, Szu-Chi Chen, Bo-Han Feng, Yun-Man Hsu, Hsiang Hsieh, Yu-Jung Lin, Yue-Ling Wu, Jia-Kai Dong, An-Yu Cheng, Yu-Han Huang, Lok-Lam Ieong, Kuan-Yu Chen, Ming-Douo Tchouang, Shao-Hua Sun, Che Lin, Jian-Jiun Ding, Hung-yi Lee
arXiv:2602. 11790v2 Announce Type: replace Abstract: Although recent end-to-end video generation models demonstrate impressive performance in visually oriented content creation, they remain limited in scenarios that require strict logical rigor and precise knowledge representation, such as instructional and educational media.
By Lingyong Yan, Jiulong Wu, Dong Xie, Weixian Shi, Deguo Xia, Jizhou Huang
SkillPE is a prompt‑engineering framework that evolves reusable cinematic skills from expert‑authored seeds to improve text‑to‑video generation for non‑experts. It encodes shot logic, composition, lighting, sound design, and other filmmaking cues in a fine‑grained format, and uses movie references classified as resonators, dissonants, and divergents to refine skill application and inspire creative alternatives. Experiments on StoryEval and VBench demonstrate up to 1.40‑point gains over the strongest baseline and 0.51 points over seed skills on a 7‑point four‑dimensional evaluation, while remaining competitive on benchmark‑native metrics.
arXiv:2607. 18529v1 Announce Type: cross Abstract: Teaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of their pedagogical quality.
By Jia-Kai Dong, Yi-Cheng Lin, Hung-yi Lee
KnowVis is a framework that converts linear video lectures into knowledge‑centric visual summaries. It first builds a detailed concept map from multimodal video content to identify key and challenging concepts, then organizes these into structured knowledge units before synthesizing engaging visual narratives. The authors also provide a dataset of 125 educational videos with 1,079 visual summaries and show through automated metrics and a human study that KnowVis outperforms existing methods in accuracy, clarity, and learning outcomes.
arXiv:2606.08091v2 Announce Type: replace
Abstract: Agentic long video generation requires planning, tool orchestration, and cross-clip coordination over a long horizon. Most existing video agents ei...
By Jianhui Wei, Yan Zhang, Jie Tan, Hengchuan Zhu, Xiaotian Zhang, Ziyi Chen, Daoan Zhang, Wei Xu, Yeying Jin, Zuozhu Liu
AI video generation has advanced rapidly and entered widespread commercial use. As a result, quality differences among videos produced by state-of-the-art AI video generation models~(AIVGMs) have become increasingly difficult to discern using conventional evaluation criteria, such as visual fidelity and semantic instruction following.
arXiv:2606. 03288v1 Announce Type: cross Abstract: Introductory programming (CS1) courses often struggle to support students' understanding of program execution.
By Yuri Noviello, Naaz Sibia, Anastasiia Birillo, Thomas Overklift Vaupel Klein, Michael Liut, Gosia Migut
arXiv:2609.37950v1 Announce Type: new
Abstract: Video understanding agents acquire evidence through an executable harness that controls what they observe and how they use those observations. However,...
By Bingjun Luo, Jialin Guo, Siqi Li
KnowVis is a framework that converts linear video lectures into knowledge‑centric visual narratives. It first extracts a detailed concept map from multimodal video content to identify key and challenging concepts, then builds structured knowledge units and synthesizes engaging visual summaries. The authors also provide a curated dataset of 125 educational videos across 10 disciplines, paired with 1,079 visual summaries, and show through automated evaluations and a human study that KnowVis produces more accurate, clear visuals that reduce cognitive load and improve learning effectiveness and knowledge retention.
By Yi Xu, Yifan Hou, Xiaoyu Zhang
arXiv:2606. 01285v1 Announce Type: cross Abstract: Text-to-video generation has advanced rapidly in visual quality, but remains under-evaluated for factuality and practical usefulness.
By Chenxu Wang, Mingda Chen