The paper investigates how changes in evaluation settings—such as camera angles and caption wording—affect the rankings of text-to-3D generators. Using 300 fixed scenes and varying eight render and caption factors, the authors find that configuration variance often exceeds generator variance, leading to frequent shifts in the top-scoring model across 19 alignment evaluators. They conclude that observed winner changes are descriptive rather than definitive, and recommend detailed reporting of generator, score, and protocol specifics to account for uncertainty.
By Anson Y. Lam, Shuqing Li, Michael R. Lyu
arXiv:2509.01167v3 Announce Type: replace-cross
Abstract: Vision-language models (VLMs) can ingest only a limited number of video frames, making frame selection a practical necessity. But do current...
By Hyunjong Ok, Jaeho Lee
A score on a temporal video question answering benchmark is meant to measure that a model has temporal understanding, but it conflates two questions. 1.
arXiv:2607. 12304v1 Announce Type: cross Abstract: A score on a temporal video question answering benchmark is meant to measure that a model has temporal understanding, but it conflates two questions.
By Farrukh Rahman
arXiv:2602.10639v2 Announce Type: replace
Abstract: Video Large Language Models (VideoLLMs) have achieved strong performance on video understanding tasks, yet existing benchmarks evaluate only what m...
By Yuxin Cao, Wei Song, Shangzhi Xu, Jingling Xue, Jin Song Dong
The study investigates how long‑video language models decide which frames to keep, compress, and reuse, testing each decision in isolation across six selection rules, three benchmarks, and two answering models. It finds that selecting frames based on queries yields the biggest performance boost, that halving spatial resolution costs little, and that reallocating saved tokens to more compressed frames can further improve accuracy. The work also highlights the importance of a unified evaluation harness to avoid misleading comparisons.
By Prakhar Khatri
arXiv:2609.13288v1 Announce Type: new
Abstract: Video-language models can answer multiple-choice questions with high confidence yet be wrong. We study whether answer-level reliability scores can be i...
By Guoxiang Ren, Rohitash Chandra
arXiv:2607. 13305v1 Announce Type: cross Abstract: Benchmark accuracy in video large language models (LLMs) is often treated as evidence of visual understanding.
By Jae Joong Lee
The paper investigates how answer candidates in masked diffusion multimodal large language models (MLLMs) can stabilize before their rationales are fully generated. It distinguishes between retrospective stabilization of the logged candidate and token commitment, and evaluates these two temporal markers across three visual question‑answering benchmarks. The study finds that a large portion of the rationale canvas remains unwritten at stabilization, that reducing block length dramatically lowers this fraction, and that direct prompting instructions can significantly alter accuracy depending on the model and dataset. The authors also use matched‑canvas image ablations to separate visual sensitivity from answer stabilization, concluding that coverage rather than conditional accuracy drives most prompting differences.
By Keuntae Kim, Yong Suk Choi
PACE (Precise AI Cinematic Expression) is a typed specification that captures a film’s spatial plan—screenplay evidence, characters, props, locations, and camera actions—at script, scene, shot, or panel levels, with inheritance to lower levels. A compiler transforms this plan into prompts for diffusion models and a metrically accurate 3D scene, while a camera solver ensures the declared framing is achieved. Evaluation on the Automatic Drive screenplay and 204 external shots shows high positional accuracy and improved action capture when poses are declared, though transitions and motion remain future work.
By Bing Duan, Qiang Guo, Linpu Li, Zhijian Mao, Min Zhu, Zhirui Ren, Yiwei Yan, Xi Chu, Xiaoding Li
arXiv:2606. 26079v1 Announce Type: cross Abstract: Standard benchmarks for multimodal large language models (MLLMs) score each item on one canonical ordering and miss whether order-irrelevant shuffling changes the answer, a baseline reliability property called for by emerging AI evaluation guidelines.
By Akshay Paruchuri, Sanmi Koyejo, Ehsan Adeli
The paper introduces Behavior Pack Optimization (BPO), a post‑training method for video multimodal large language models that replaces single‑response rewards with a set of outputs across counterfactual views. BPO enforces stability when interventions are irrelevant, sensitivity when key evidence is removed, and abstention when no evidence remains, using an anchor‑relative advantage to keep the objective stable with small pack sizes. Experiments on datasets such as TempCompass, MVBench, and NExT‑QA show that BPO improves macro accuracy and abstention metrics for models like Qwen2.5‑VL‑7B‑Instruct, with gains that transfer to other benchmarks and models.
By Zhaolu Kang, Shiyu Liu, Tailong Luo, Wei Zhang, Yingjie He, Lei Wei, Guansu Wang, Liang He, Siheng Wang, Guangyuan Dong, Jiaqi Su, Shuang Chen, Haoyu Ji, Qishi Zhan, Kaiyue Zhou