arXiv Computer Vision

Stable Scores, Unstable Answers: Frame Phase and Option Order in Video Multiple-Choice Evaluation

arXiv AI
6d ago

Frozen Scenes, Shifting Winners: Configuration Fragility in Text-to-3D Evaluation

The paper investigates how changes in evaluation settings—such as camera angles and caption wording—affect the rankings of text-to-3D generators. Using 300 fixed scenes and varying eight render and caption factors, the authors find that configuration variance often exceeds generator variance, leading to frequent shifts in the top-scoring model across 19 alignment evaluators. They conclude that observed winner changes are descriptive rather than definitive, and recommend detailed reporting of generator, score, and protocol specifics to account for uncertainty.

By Anson Y. Lam, Shuqing Li, Michael R. Lyu
arXiv Computation and Language
Sep 4

Select, Compress, Reinvest: A Controlled Study of Visual-Token Allocation in Long-Video MLLMs

The study investigates how long‑video language models decide which frames to keep, compress, and reuse, testing each decision in isolation across six selection rules, three benchmarks, and two answering models. It finds that selecting frames based on queries yields the biggest performance boost, that halving spatial resolution costs little, and that reallocating saved tokens to more compressed frames can further improve accuracy. The work also highlights the importance of a unified evaluation harness to avoid misleading comparisons.

By Prakhar Khatri
arXiv AI
6d ago

Two Clocks in Diffusion MLLMs: When Answers Stabilize Before Rationales Unfold

The paper investigates how answer candidates in masked diffusion multimodal large language models (MLLMs) can stabilize before their rationales are fully generated. It distinguishes between retrospective stabilization of the logged candidate and token commitment, and evaluates these two temporal markers across three visual question‑answering benchmarks. The study finds that a large portion of the rationale canvas remains unwritten at stabilization, that reducing block length dramatically lowers this fraction, and that direct prompting instructions can significantly alter accuracy depending on the model and dataset. The authors also use matched‑canvas image ablations to separate visual sensitivity from answer stabilization, concluding that coverage rather than conditional accuracy drives most prompting differences.

By Keuntae Kim, Yong Suk Choi
arXiv AI
Sep 18

PACE: Precise AI Cinematic Expression: A Typed Specification for Script-Grounded Previsualization and Geometric Conformance

PACE (Precise AI Cinematic Expression) is a typed specification that captures a film’s spatial plan—screenplay evidence, characters, props, locations, and camera actions—at script, scene, shot, or panel levels, with inheritance to lower levels. A compiler transforms this plan into prompts for diffusion models and a metrically accurate 3D scene, while a camera solver ensures the declared framing is achieved. Evaluation on the Automatic Drive screenplay and 204 external shots shows high positional accuracy and improved action capture when poses are declared, though transitions and motion remain future work.

By Bing Duan, Qiang Guo, Linpu Li, Zhijian Mao, Min Zhu, Zhirui Ren, Yiwei Yan, Xi Chu, Xiaoding Li
arXiv Machine Learning
Jun 25

Same Evidence, Different Answer: Auditing Order Sensitivity in Multimodal Large Language Models

arXiv:2606. 26079v1 Announce Type: cross Abstract: Standard benchmarks for multimodal large language models (MLLMs) score each item on one canonical ordering and miss whether order-irrelevant shuffling changes the answer, a baseline reliability property called for by emerging AI evaluation guidelines.

By Akshay Paruchuri, Sanmi Koyejo, Ehsan Adeli
arXiv Computer Vision
3d ago

Behavior Pack Optimization for Video MLLM Post-Training

The paper introduces Behavior Pack Optimization (BPO), a post‑training method for video multimodal large language models that replaces single‑response rewards with a set of outputs across counterfactual views. BPO enforces stability when interventions are irrelevant, sensitivity when key evidence is removed, and abstention when no evidence remains, using an anchor‑relative advantage to keep the objective stable with small pack sizes. Experiments on datasets such as TempCompass, MVBench, and NExT‑QA show that BPO improves macro accuracy and abstention metrics for models like Qwen2.5‑VL‑7B‑Instruct, with gains that transfer to other benchmarks and models.

By Zhaolu Kang, Shiyu Liu, Tailong Luo, Wei Zhang, Yingjie He, Lei Wei, Guansu Wang, Liang He, Siheng Wang, Guangyuan Dong, Jiaqi Su, Shuang Chen, Haoyu Ji, Qishi Zhan, Kaiyue Zhou