arXiv Computer Vision By Andrew Franck, Brendan Ng, Ben Fitzgerald, Zane Derrod, Chris Cianci, Chris Craney

VISTA: Dense Multi-Label Classroom Coding with Vision-Language Models

Read the original on arXiv Computer Vision →

The paper introduces VISTA, a baseline for dense multi‑label classroom coding that leverages the COPUS protocol as a video benchmark. VISTA applies MiniCPM‑V‑4.5 over sliding windows, refines predictions with an MLP head, and aggregates results onto a 2‑minute COPUS grid, achieving 80.1% macro accuracy on held‑out chemistry lectures. The authors also identify systematic failure modes and provide benchmark tooling and code on GitHub.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
Aug 27

Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios

Video-IFBench is a new benchmark designed to evaluate how well multimodal large language models (MLLMs) follow user-specified instructions in video understanding tasks. It introduces an instruction taxonomy with four templates—single-task, multi-task, selection, and nested—covering 32 task types and 39 constraint categories that span semantic and format requirements. The benchmark was built using a semi-automatic pipeline that combines MLLMs, programmatic processing, and human verification, producing 1.5K samples, and a large-scale evaluation of over 20 recent MLLMs shows that instruction following remains difficult, especially for complex constraints and conditional structures.

By Hongbo Liu, Peixian Chen, Sihan Liu, Peiyuan Zhang, Kai Zou, Dian Zheng, Xiaoxing Hu, Yuhao Dong, Mengdan Zhang, Yunhang Shen, Haoyu Cao, Wei Liu, Weibo Gu, Xing Sun, Shengjie Zhao
Hugging Face Trending Papers
Jun 20

Zero-Shot Vision-Language Models for Classroom Engagement Recognition: A Benchmark Study of Prompt Sensitivity and Cross-Dataset Generalization

Automated classroom engagement recognition holds substantial promise for scalable learning analytics, yet the suitability of modern Vision-Language Models (VLMs) for this task under zero-shot conditions remains largely unexplored. We present a systematic benchmark that evaluates five widely-used VLMs: CLIP, BLIP-VQA, GPT-4o, LLaVA-1.

arXiv AI
Aug 14

EgoMonth: A Month-Level Egocentric Video Benchmark for Long-Term Spatiotemporal Memory

arXiv:2608. 13113v1 Announce Type: cross Abstract: Recent advances in Multimodal Large Language Models (MLLMs) have led to substantial progress in video understanding, accompanied by a growing number of long video benchmarks.

By Weitao Chen, Hu Jiaxin, Xie Tianyidan, Yang Li, Yuyi Qian, Banghao Xu, Ziheng Tang, Shenyi Wang, Mingyue Yu, Duo Li, Jiacheng Shi, Gao Wang, Zhan Xu, Zhicheng Qiu, Xuanfu Li, Jian Yang, Lanjun Wang, Zili Yi
arXiv AI
2d ago

SONIC-O1: A Real-World Benchmark for Evaluating Multimodal Large Language Models on Audio-Video Understanding

SONIC‑O1 is a new benchmark designed to evaluate multimodal large language models on audio‑video understanding. It contains 60 hours of 231 clips across 13 real‑world conversational domains, with 4,958 human‑verified annotations and demographic metadata. The benchmark tests open‑ended summarization, multiple‑choice question answering, and temporally grounded reasoning, revealing performance gaps between model families and across demographic groups.

By Ahmed Y. Radwan, Christos Emmanouilidis, Hina Tabassum, Deval Pandya, Shaina Raza