arXiv AI

Dynamic Learning Solutions: A System for Personalized Educational Video Generation

Dynamic Learning Solutions is an automated pipeline that transforms NCERT textbook PDFs into interactive video explanations. Users upload a PDF and ask a question; the system retrieves relevant content, generates a multi-scene script, creates images with Stable Diffusion, animates them with DynamiCrafter, and adds synchronized narration via Google Text‑to‑Speech. The result is a coherent, textbook‑aligned video that turns static material into an engaging learning experience.

arXiv Computation and Language
Sep 4

KnowVis: Knowledge-Centric Visual Summarization for Video Lectures

KnowVis is a framework that converts linear video lectures into knowledge‑centric visual narratives. It first extracts a detailed concept map from multimodal video content to identify key and challenging concepts, then builds structured knowledge units and synthesizes engaging visual summaries. The authors also provide a curated dataset of 125 educational videos across 10 disciplines, paired with 1,079 visual summaries, and show through automated evaluations and a human study that KnowVis produces more accurate, clear visuals that reduce cognitive load and improve learning effectiveness and knowledge retention.

By Yi Xu, Yifan Hou, Xiaoyu Zhang
arXiv AI
Jun 2

Beyond End-to-End Video Models: An LLM-Based Multi-Agent System for Educational Video Generation

arXiv:2602. 11790v2 Announce Type: replace Abstract: Although recent end-to-end video generation models demonstrate impressive performance in visually oriented content creation, they remain limited in scenarios that require strict logical rigor and precise knowledge representation, such as instructional and educational media.

By Lingyong Yan, Jiulong Wu, Dong Xie, Weixian Shi, Deguo Xia, Jizhou Huang
Hugging Face Trending Papers
Sep 3

KnowVis: Knowledge-Centric Visual Summarization for Video Lectures

KnowVis is a framework that converts linear video lectures into knowledge‑centric visual summaries. It first builds a detailed concept map from multimodal video content to identify key and challenging concepts, then organizes these into structured knowledge units before synthesizing engaging visual narratives. The authors also provide a dataset of 125 educational videos with 1,079 visual summaries and show through automated metrics and a human study that KnowVis outperforms existing methods in accuracy, clarity, and learning outcomes.

arXiv AI
Aug 28

Omni-Interactive Universal Embedder

The paper introduces Omni-Interactive Universal Embedder (OmniUE), a unified embedding framework that learns a single representation space for text, video, and audio using learnable tokens and intermediate-layer representations. OmniUE supports omni-interactive querying, allowing users to input text, visual regions, or audio spans, which are processed by segmenters and an omni-LLM to generate user-conditioned embeddings. The authors evaluate OmniUE on the new OmniCHOIR benchmark and other multimodal tasks, reporting significant performance gains over state‑of‑the‑art baselines across textual, audio, and visual interactive settings.

By Wei-Yao Wang, Kazuya Tateishi, Shuyang Cui, Christian Simon, Takashi Shibuya, Shusuke Takahashi, Yuki Mitsufuji