arXiv Computation and Language By Prakhar Khatri

Select, Compress, Reinvest: A Controlled Study of Visual-Token Allocation in Long-Video MLLMs

Read the original on arXiv Computation and Language →

The study investigates how long‑video language models decide which frames to keep, compress, and reuse, testing each decision in isolation across six selection rules, three benchmarks, and two answering models. It finds that selecting frames based on queries yields the biggest performance boost, that halving spatial resolution costs little, and that reallocating saved tokens to more compressed frames can further improve accuracy. The work also highlights the importance of a unified evaluation harness to avoid misleading comparisons.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computer Vision
Sep 10

MotionBlind: Probing the Illusion of Motion Understanding in Video-LLMs

arXiv:2609.09528v1 Announce Type: new Abstract: Video large language models (Video-LLMs) are increasingly used as the perceptual front end of world models, a role that assumes they can read motion: h...

By Dhairya Bhatia, Bishoy Galoaa, Oliver Fritsche, Shahid Kamal, Muhammad Obaidullah Abdul Salam, Umer Saleem, Om Rastogi, Frania Felix Chettiar, Nesli Erdogmus, Sarah Ostadabbas