arXiv AI By Fan Yuxuan, Huang Miaojun, Zhang Haimei, Wu Jingshen, Liu Hao

PhoenixNest-Video: Evidence-Grounded Multimodal Agent Framework for Automated Video Interview Assessment

Read the original on arXiv AI →

PhoenixNest-Video is an evidence‑grounded multimodal agent designed for automated video interview assessment. It constructs a semantic video graph as working memory, retrieves information conditioned on rubrics across visual, audio, and textual streams, and outputs per‑criterion scores tied to the candidate’s materials. Trained with rubric‑based reinforcement learning, the system achieves 91.50% grade‑level accuracy on VInterview‑2025, outperforming larger proprietary models while providing traceable evidence for each score.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 12

Frozen Multimodal Embeddings for AI-Assisted Interview Assessment of Personality and Cognitive Ability

arXiv:2606. 11930v2 Announce Type: replace-cross Abstract: Predicting psychological traits from asynchronous video interviews (AVIs) is a challenging problem in AI-assisted interview assessment because labeled datasets are limited while each response contains high-dimensional visual, acoustic, and verbal signals.

By Kuo-En Hung, Hung-Yue Suen, Shih-Ching Yeh, Hsiang-Wen Wang
arXiv Machine Learning
Jul 7

Incentivizing Vision Language Models to Search for Long Video Question Answering

arXiv:2607. 02959v1 Announce Type: cross Abstract: We introduce VSeek, an agentic framework that transforms long-video question answering (LVQA) from a passive, single-pass perception task into a multi-turn retrieval process.

By Harsh Goel, S P Sharan, Sahil Shah, Minkyu Choi, Joungbin An, Kristen Grauman, Sandeep P. Chinchali
arXiv AI
Aug 28

Beyond Factual QA: Mentorship-Oriented Question Answering over Long-Form Multilingual Content

The paper introduces MentorQA, a multilingual dataset and evaluation framework for mentorship-oriented question answering derived from long‑form videos. It contains nearly 9,000 QA pairs across four languages and defines evaluation dimensions such as clarity, alignment, and learning value that extend beyond factual accuracy. Experiments show that Multi‑Agent QA pipelines outperform other architectures, especially on complex topics and low‑resource languages, while automated LLM‑based evaluation shows variable alignment with human judgments.

By Parth Bhalerao, Diola Dsouza, Ruiwen Guan, Oana Ignat
arXiv Computer Vision
Aug 25

FIRM-Video: Check Before You Score for Reliable Text-to-Video Reward Modeling

arXiv:2608.21839v1 Announce Type: new Abstract: Reliable reward models are essential for text-to-video evaluation and alignment. However, the trade-off between evaluation accuracy and inference effic...

By Peiyuan Zhang, Xiangyu Zhao, Hongbo Liu, Xiaoxing Hu, Mingxin Liu, Shuran Ma, Yunhang Shen, Jian Hu, Haihan Gao, Haoyu Cao, Xue Yang
arXiv AI
Jun 11

Frozen Multimodal Embeddings for Personality and Cognitive Ability Assessment in Asynchronous Video Interviews

arXiv:2606. 11930v1 Announce Type: cross Abstract: Predicting psychological traits from asynchronous video interviews (AVIs) is a challenging multimodal learning problem because labeled datasets are limited while each response contains high-dimensional visual, acoustic, and verbal signals.

By Kuo-En Hung, Hung-Yue Suen, Shih-Ching Yeh, Hsiang-Wen Wang