KnowVis is a framework that converts linear video lectures into knowledge‑centric visual narratives. It first extracts a detailed concept map from multimodal video content to identify key and challenging concepts, then builds structured knowledge units and synthesizes engaging visual summaries. The authors also provide a curated dataset of 125 educational videos across 10 disciplines, paired with 1,079 visual summaries, and show through automated evaluations and a human study that KnowVis produces more accurate, clear visuals that reduce cognitive load and improve learning effectiveness and knowledge retention.
By Yi Xu, Yifan Hou, Xiaoyu Zhang
arXiv:2606. 01285v1 Announce Type: cross Abstract: Text-to-video generation has advanced rapidly in visual quality, but remains under-evaluated for factuality and practical usefulness.
By Chenxu Wang, Mingda Chen
arXiv:2609.24083v1 Announce Type: new
Abstract: Generative AI enables scalable production of educational videos, but current systems largely focus on producing visually coherent content rather than s...
By Xinchen Ma, Shuimu Wang, Gaole He, Yanbin Zhang, Chunyang Wang, Yunshi Lan, Weining Qian
arXiv:2606. 14762v1 Announce Type: cross Abstract: As video content continues to expand across educational platforms, recorded lectures, and live-streamed entertainment, the need for efficient and structured analysis of long-form footage has increased \cite{1}.
By Julian Abelarde, Hugo Garrido-Lestache Belinchon
arXiv:2608. 03161v1 Announce Type: new Abstract: Lecture videos distribute knowledge across speech, slide text, diagrams, equations, and presentation order, which transcript-only retrieval does not fully preserve.
By Sahil Al Farib, Momota Ahsana Meem, Sheikh Redwanul Islam, Md. Tanvir Raihan
The paper introduces Clue-OPSD, a clue‑privileged on‑policy self‑distillation framework that improves long‑video understanding by focusing on short, question‑relevant clue intervals rather than the entire video. Experiments on multiple benchmarks and Qwen3.5 model scales show that this approach consistently outperforms standard backbone models and competes strongly with supervised post‑training baselines, all while requiring fewer input frames and no additional inference modules.
By Kaishen Wang, Dongdi Zhao, Yijun Liang, Dingqiang Ye, Ruibo Chen, Heng Huang, Di Fu
The paper introduces Concept-Driven Domain Adaptation (CDDA), a three-stage framework that adapts vision‑language models for concept‑to‑example video retrieval in educational settings. CDDA first structures textual embeddings using textbook and teacher‑handbook concept pairs, then transfers this geometry to documentary visuals with a frozen visual encoder, and finally jointly fine‑tunes both encoders with sparse visual concept supervision. On a middle‑school physics benchmark, CDDA outperforms several multimodal baselines in retrieving concept‑driven moments while preserving concrete image‑text alignment.
By Haiming Zhao, Tai Wang, Kun Zhang, Xicheng Peng, Zhiyang Li
arXiv:2607. 17994v1 Announce Type: cross Abstract: Video understanding has become more and more important with the growth of Artificial Intelligence (AI) for video generation.
By Rui Chu, Yingjie Lao
Vision-language models (VLMs) have made substantial progress in long-video understanding, with standard backbone models typically answering questions from frames sampled across the full video. However...
arXiv:2604.17422v2 Announce Type: replace
Abstract: Long video understanding remains a formidable challenge for Multimodal Large Language Models (MLLMs) due to the prohibitive cost of processing dens...
By Shaoguang Wang, Weiyu Guo, Ziyang Chen, Xuming Hu, Hui Xiong
arXiv:2606. 07433v1 Announce Type: cross Abstract: Video understanding is being rapidly transformed by multimodal large language models (MLLMs), as research moves from short clips to long, multimodal, and knowledge-intensive video scenarios.
By Jiahao Meng, Yue Tan, Qi Xu, Kuan Gao, Weisong Liu, Yanwei Li, Jason Li, Lingdong Kong, Haochen Wang, Qianyu Zhou, Jiangning Zhang, Guangliang Cheng, Yunhai Tong, Lu Qi, Minghsuan Yang
Dynamic Learning Solutions is an automated pipeline that transforms NCERT textbook PDFs into interactive video explanations. Users upload a PDF and ask a question; the system retrieves relevant content, generates a multi-scene script, creates images with Stable Diffusion, animates them with DynamiCrafter, and adds synchronized narration via Google Text‑to‑Speech. The result is a coherent, textbook‑aligned video that turns static material into an engaging learning experience.
By Siddhanth Sridhar, Shreya Chaurasia, Baddela Sai Yaswantha Reddy, Deepak Parmar, Shylaja S S