Automated classroom engagement recognition holds substantial promise for scalable learning analytics, yet the suitability of modern Vision-Language Models (VLMs) for this task under zero-shot conditions remains largely unexplored. We present a systematic benchmark that evaluates five widely-used VLMs: CLIP, BLIP-VQA, GPT-4o, LLaVA-1.
arXiv:2609.25001v1 Announce Type: new
Abstract: Modern video games provide a measurable testbed for AI models, combining abilities of visual understanding, instruction decomposition, goal planning, a...
By Yiran Wang, Xingyilang Yin, Junfu Pu, Guangzhi Wang, Kaifeng Li, Mingyu Ouyang, Huiqiang Sun, Lingen Li, Cheng Cheng, Wangbo Yu, Honghao Chen, Xiaodong Cun, Chi-Man Pun, Zhiguo Cao, Ying Shan
arXiv:2605. 04733v2 Announce Type: replace Abstract: Text-based role-playing models can imitate character styles, but often fail to capture scene atmosphere and evolving tension, which are crucial for immersive applications such as VR games and interactive narratives.
By Miao Wang, Yuling Shi, Yijiang Li, Yeheng Chen, Xiaodong Gu, Bin Li, Bo Gao, Jun Wang, Zengxin Han, Jingtong Wu, Yaduan Ruan
arXiv:2609.05517v1 Announce Type: cross
Abstract: Human observers prioritize visual information according to task goals. Most computational models of naturalistic viewing are gaze-trained for free vi...
By Han Zhang
arXiv:2606. 00616v1 Announce Type: cross Abstract: Recent Vision-Language Models (VLMs) struggle with grounded reasoning, temporal consistency, and context aware planning in videos.
By Shivam Singh, Saptarshi Majumdar, Pratik Prabhanjan, Zicheng Liu, Emad Barsoum
arXiv:2606. 12550v1 Announce Type: cross Abstract: Open-world mapless navigation from sparse language instructions requires resolving underspecified goals and inferring which environmental cues are relevant for reaching the goal.
By Arthur Zhang, Carl Qi, Donne Su, Xiangyun Meng, Amy Zhang, Joydeep Biswas
SocialReasonBench is a new video‑multiple‑choice QA benchmark designed to test socially grounded reasoning in interactive narrative videos. It uses branching gameplay footage from *Detroit: Become Human*, where player choices create alternative social outcomes that can be verified against the game’s script and flowchart. The benchmark includes seven reasoning dimensions—such as intent recognition, emotional empathy, moral dilemma, counterfactual reasoning, and causal antecedent—and employs a multi‑agent pipeline to curate clips, ground answer labels, and generate theory‑guided questions with diagnostic distractors.
By Zheyu Huang, Zijing Shi, Haozhe Luo, Huadong Tang, Mingyu Liu, Meng Fang, Ling Chen
Hidden‑Shot introduces an implicit prompt mechanism that extracts task‑specific visual information and merges it with in‑task processing to boost one‑shot performance on new low‑level vision tasks. The method injects this prompt cost‑effectively while minimally altering the base generalist model’s architecture. A data‑driven evaluation framework, C/U assessment, is proposed to systematically test generalization across conventional and unconventional tasks, and experiments on seven and ten datasets show Hidden‑Shot outperforming state‑of‑the‑art models.
By Shao-Jun Xia, Xianzheng Ma, Zichong Meng
arXiv:2511. 18735v3 Announce Type: replace-cross Abstract: In this work, we define Foresight Intelligence as the capability to anticipate and interpret future events-an ability essential for applications such as autonomous driving, yet largely overlooked by existing research.
By Zhantao Gong, Liaoyuan Fan, Qing Guo, Xun Xu, Xulei Yang, Shijie Li
The paper investigates whether Vision‑Language Models (VLMs) can understand visual persuasiveness by evaluating image‑message pairs that humans consistently judge as persuasive. It introduces Visual Persuasive Factors (VPFs), a taxonomy from cognitive psychology, to quantify visual cues influencing persuasive judgments. Empirical analysis shows VLMs tend to over‑predict persuasiveness, partially reproducing human patterns but often generating false positives, and that VPF‑guided interventions can improve performance only when properly framed.
By Gyuwon Park, Hyounghun Kim
arXiv:2609.31665v2 Announce Type: replace
Abstract: Action anticipation predicts future human actions from partial video under incomplete context and temporal uncertainty. Recent systems introduce la...
By Mahsa Mohammadi, Zeyu Fu, Sareh Rowlands
The paper introduces VISTA, a visual harness that equips a general-purpose multimodal model with long‑horizon vision and a lossless visual memory. VISTA enables the model to directly perceive and actively retrieve past observations, allowing it to reorganize visual input during reasoning. On the ARC‑AGI‑3 benchmark, VISTA boosts Claude Opus 5.0’s Relative Human Action Efficiency from 40.68 to a perfect 100.00, completing all 25 public games with 57.4% fewer actions than first‑time human participants, and it also outperforms baselines on three additional visual game and puzzle benchmarks.
By Qiushi Han, Keya Hu, Linlu Qiu, Cathy Wu, Kaiming He