arXiv AI
6d ago

Frozen Scenes, Shifting Winners: Configuration Fragility in Text-to-3D Evaluation

The paper investigates how changes in evaluation settings—such as camera angles and caption wording—affect the rankings of text-to-3D generators. Using 300 fixed scenes and varying eight render and caption factors, the authors find that configuration variance often exceeds generator variance, leading to frequent shifts in the top-scoring model across 19 alignment evaluators. They conclude that observed winner changes are descriptive rather than definitive, and recommend detailed reporting of generator, score, and protocol specifics to account for uncertainty.

By Anson Y. Lam, Shuqing Li, Michael R. Lyu
arXiv Computation and Language
Sep 4

Select, Compress, Reinvest: A Controlled Study of Visual-Token Allocation in Long-Video MLLMs

The study investigates how long‑video language models decide which frames to keep, compress, and reuse, testing each decision in isolation across six selection rules, three benchmarks, and two answering models. It finds that selecting frames based on queries yields the biggest performance boost, that halving spatial resolution costs little, and that reallocating saved tokens to more compressed frames can further improve accuracy. The work also highlights the importance of a unified evaluation harness to avoid misleading comparisons.

By Prakhar Khatri