KnowVis is a framework that converts linear video lectures into knowledge‑centric visual summaries. It first builds a detailed concept map from multimodal video content to identify key and challenging concepts, then organizes these into structured knowledge units before synthesizing engaging visual narratives. The authors also provide a dataset of 125 educational videos with 1,079 visual summaries and show through automated metrics and a human study that KnowVis outperforms existing methods in accuracy, clarity, and learning outcomes.
arXiv:2606. 14762v1 Announce Type: cross Abstract: As video content continues to expand across educational platforms, recorded lectures, and live-streamed entertainment, the need for efficient and structured analysis of long-form footage has increased \cite{1}.
By Julian Abelarde, Hugo Garrido-Lestache Belinchon
arXiv:2606. 01285v1 Announce Type: cross Abstract: Text-to-video generation has advanced rapidly in visual quality, but remains under-evaluated for factuality and practical usefulness.
By Chenxu Wang, Mingda Chen
The paper introduces Concept-Driven Domain Adaptation (CDDA), a three-stage framework that adapts vision‑language models for concept‑to‑example video retrieval in educational settings. CDDA first structures textual embeddings using textbook and teacher‑handbook concept pairs, then transfers this geometry to documentary visuals with a frozen visual encoder, and finally jointly fine‑tunes both encoders with sparse visual concept supervision. On a middle‑school physics benchmark, CDDA outperforms several multimodal baselines in retrieving concept‑driven moments while preserving concrete image‑text alignment.
By Haiming Zhao, Tai Wang, Kun Zhang, Xicheng Peng, Zhiyang Li
arXiv:2609.24083v1 Announce Type: new
Abstract: Generative AI enables scalable production of educational videos, but current systems largely focus on producing visually coherent content rather than s...
By Xinchen Ma, Shuimu Wang, Gaole He, Yanbin Zhang, Chunyang Wang, Yunshi Lan, Weining Qian
arXiv:2608. 03161v1 Announce Type: new Abstract: Lecture videos distribute knowledge across speech, slide text, diagrams, equations, and presentation order, which transcript-only retrieval does not fully preserve.
By Sahil Al Farib, Momota Ahsana Meem, Sheikh Redwanul Islam, Md. Tanvir Raihan
The paper introduces Clue-OPSD, a clue‑privileged on‑policy self‑distillation framework that improves long‑video understanding by focusing on short, question‑relevant clue intervals rather than the entire video. Experiments on multiple benchmarks and Qwen3.5 model scales show that this approach consistently outperforms standard backbone models and competes strongly with supervised post‑training baselines, all while requiring fewer input frames and no additional inference modules.
By Kaishen Wang, Dongdi Zhao, Yijun Liang, Dingqiang Ye, Ruibo Chen, Heng Huang, Di Fu
arXiv:2607. 17994v1 Announce Type: cross Abstract: Video understanding has become more and more important with the growth of Artificial Intelligence (AI) for video generation.
By Rui Chu, Yingjie Lao
arXiv:2604.17422v2 Announce Type: replace
Abstract: Long video understanding remains a formidable challenge for Multimodal Large Language Models (MLLMs) due to the prohibitive cost of processing dens...
By Shaoguang Wang, Weiyu Guo, Ziyang Chen, Xuming Hu, Hui Xiong
arXiv:2606. 07433v1 Announce Type: cross Abstract: Video understanding is being rapidly transformed by multimodal large language models (MLLMs), as research moves from short clips to long, multimodal, and knowledge-intensive video scenarios.
By Jiahao Meng, Yue Tan, Qi Xu, Kuan Gao, Weisong Liu, Yanwei Li, Jason Li, Lingdong Kong, Haochen Wang, Qianyu Zhou, Jiangning Zhang, Guangliang Cheng, Yunhai Tong, Lu Qi, Minghsuan Yang
arXiv:2604. 10024v2 Announce Type: replace-cross Abstract: Long video summarization presents significant challenges for multimodal large language models (MLLMs), particularly in maintaining temporal fidelity over extended durations and producing summaries that are both semantically and temporally grounded.
By Alkesh Patel, Melis Ozyildirim, Ying-Chang Cheng, Ganesh Nagarajan
arXiv:2607. 06618v1 Announce Type: cross Abstract: Following the CMIVQA, MMI-VQA, and M4IVQA challenges in NLPCC 2023--2025, we introduce the Difficulty-Aware Medical Instructional Video Question Answering (DA-MIVQA) shared task for NLPCC 2026.
By Shenxi Liu, Kan Li, Mingyang Zhao, Yuhang Tian, Bin Li