arXiv AI By Jia-Kai Dong, Yi-Cheng Lin, Hung-yi Lee

EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration

Read the original on arXiv AI →

arXiv:2607. 18529v1 Announce Type: cross Abstract: Teaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of their pedagogical quality.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 11

Findings of the First Teaching Monster Challenge: A Benchmark of Pedagogical Content Knowledge in AI Agents

arXiv:2608. 08852v1 Announce Type: new Abstract: AI agents can now solve problems, answer like subject experts, and generate long-form multimodal content.

By Yi-Cheng Lin, Yu-Kai Guo, Szu-Chi Chen, Bo-Han Feng, Yun-Man Hsu, Hsiang Hsieh, Yu-Jung Lin, Yue-Ling Wu, Jia-Kai Dong, An-Yu Cheng, Yu-Han Huang, Lok-Lam Ieong, Kuan-Yu Chen, Ming-Douo Tchouang, Shao-Hua Sun, Che Lin, Jian-Jiun Ding, Hung-yi Lee
arXiv AI
Sep 3

PhoenixNest-Video: Evidence-Grounded Multimodal Agent Framework for Automated Video Interview Assessment

PhoenixNest-Video is an evidence‑grounded multimodal agent designed for automated video interview assessment. It constructs a semantic video graph as working memory, retrieves information conditioned on rubrics across visual, audio, and textual streams, and outputs per‑criterion scores tied to the candidate’s materials. Trained with rubric‑based reinforcement learning, the system achieves 91.50% grade‑level accuracy on VInterview‑2025, outperforming larger proprietary models while providing traceable evidence for each score.

By Fan Yuxuan, Huang Miaojun, Zhang Haimei, Wu Jingshen, Liu Hao
Hugging Face Trending Papers
Aug 10

RAVEN-Eval: Rubric-Guided Automatic Evaluation for AI Video Generation Models Based on LMM Preference Judgement

AI video generation has advanced rapidly and entered widespread commercial use. As a result, quality differences among videos produced by state-of-the-art AI video generation models~(AIVGMs) have become increasingly difficult to discern using conventional evaluation criteria, such as visual fidelity and semantic instruction following.