arXiv:2609.24083v1 Announce Type: new
Abstract: Generative AI enables scalable production of educational videos, but current systems largely focus on producing visually coherent content rather than s...
By Xinchen Ma, Shuimu Wang, Gaole He, Yanbin Zhang, Chunyang Wang, Yunshi Lan, Weining Qian
arXiv:2608. 08852v1 Announce Type: new Abstract: AI agents can now solve problems, answer like subject experts, and generate long-form multimodal content.
By Yi-Cheng Lin, Yu-Kai Guo, Szu-Chi Chen, Bo-Han Feng, Yun-Man Hsu, Hsiang Hsieh, Yu-Jung Lin, Yue-Ling Wu, Jia-Kai Dong, An-Yu Cheng, Yu-Han Huang, Lok-Lam Ieong, Kuan-Yu Chen, Ming-Douo Tchouang, Shao-Hua Sun, Che Lin, Jian-Jiun Ding, Hung-yi Lee
arXiv:2602. 11790v2 Announce Type: replace Abstract: Although recent end-to-end video generation models demonstrate impressive performance in visually oriented content creation, they remain limited in scenarios that require strict logical rigor and precise knowledge representation, such as instructional and educational media.
By Lingyong Yan, Jiulong Wu, Dong Xie, Weixian Shi, Deguo Xia, Jizhou Huang
SkillPE is a prompt‑engineering framework that evolves reusable cinematic skills from expert‑authored seeds to improve text‑to‑video generation for non‑experts. It encodes shot logic, composition, lighting, sound design, and other filmmaking cues in a fine‑grained format, and uses movie references classified as resonators, dissonants, and divergents to refine skill application and inspire creative alternatives. Experiments on StoryEval and VBench demonstrate up to 1.40‑point gains over the strongest baseline and 0.51 points over seed skills on a 7‑point four‑dimensional evaluation, while remaining competitive on benchmark‑native metrics.
arXiv:2607. 18529v1 Announce Type: cross Abstract: Teaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of their pedagogical quality.
By Jia-Kai Dong, Yi-Cheng Lin, Hung-yi Lee
KnowVis is a framework that converts linear video lectures into knowledge‑centric visual summaries. It first builds a detailed concept map from multimodal video content to identify key and challenging concepts, then organizes these into structured knowledge units before synthesizing engaging visual narratives. The authors also provide a dataset of 125 educational videos with 1,079 visual summaries and show through automated metrics and a human study that KnowVis outperforms existing methods in accuracy, clarity, and learning outcomes.