Paying for Too Many Tokens? Valid and Cost-Efficient Multimodal LLM Annotation with Simple Heuristics
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
Video provides a rich record of human behavior, interaction, and situated contexts, offering important evidence for understanding people and conducting human-centered research. As vision-language mode...
The paper investigates when vision‑language models (VLMs) can independently analyze human‑centered video and when human oversight is still needed. By reviewing 1,702 CHI 2026 papers, the authors develop a five‑dimensional taxonomy of video annotation tasks and build a benchmark of 15 representative tasks. Experiments show that VLMs alone achieve near‑human accuracy (HNS = 97.0), while human verification of VLM outputs yields the highest accuracy (HNS = 121.5) and significantly reduces annotation time and cost.
arXiv:2609.16722v1 Announce Type: new Abstract: Scaling Multimodal Large Language Models (MLLMs) to long-form video understanding is bottlenecked by the explosion of visual tokens, which saturates co...
Scaling Multimodal Large Language Models (MLLMs) to long-form video understanding is bottlenecked by the explosion of visual tokens, which saturates context windows and incurs prohibitive costs. Curre...
arXiv:2608. 09176v1 Announce Type: cross Abstract: Visual token compression for vision--language models (VLMs) has largely relied on criteria such as attention, redundancy, and uncertainty to maximize average accuracy under a fixed compute budget, implicitly assuming that all errors carry equal cost.
arXiv:2607. 06875v1 Announce Type: cross Abstract: Understanding and forecasting audience reactions to video content are crucial for improving content creation, recommendation systems, and media analysis.