FORTE: Adaptive Scoring and Exact Keyframe Selection for Long-Video Question Answering
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2607. 25266v2 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) have enabled long-form video understanding at a scale that was not previously possible.
arXiv:2609.15408v1 Announce Type: cross Abstract: Long-video understanding remains challenging for multimodal large language models (MLLMs) because densely encoding long frame sequences is computatio...
arXiv:2604.17422v2 Announce Type: replace Abstract: Long video understanding remains a formidable challenge for Multimodal Large Language Models (MLLMs) due to the prohibitive cost of processing dens...
The paper evaluates five training‑free, plug‑and‑play keyframe selection methods for multimodal large language models (MLLMs) on long‑video understanding tasks. It compares these methods across three different MLLMs and three video question‑answering benchmarks, finding that QAaF performs best in 13 of 15 settings while FOCUS ranks second. The study offers a unified benchmark for assessing MLLM‑agnostic keyframe selection techniques.
arXiv:2607. 24794v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) demonstrate superior generalization in fundamental video tasks, restricted context windows limit their long video understanding.
arXiv:2608.05707v2 Announce Type: replace Abstract: Frame selection is essential for applying Large Multimodal Models (LMMs) to long videos due to severe frame redundancy and limited context windows....