← Back to all news
arXiv Computer Vision September 3, 2026 By Shuyao Xiao, Shengling Wang, Haoyu Niu, Ke Chao, Changwei Xu, Xinran Duan, Chaoyong Jiang

Who Drives the Probability Game of VLMs? A Temporal Causal Drive Evaluation Framework

Read the original on arXiv Computer Vision →

The Flow has not summarised this story yet — read it at arXiv Computer Vision.

  • llms
  • multimodal

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

Hugging Face Trending Papers
2d ago

Who Drives the Probability Game of VLMs? A Temporal Causal Drive Evaluation Framework

Vision-language models (VLMs) are increasingly evaluated on complex image and video understanding tasks, yet conventional metrics primarily assess final-answer quality and reveal little about how diff...

llmsmultimodal
More like this →
arXiv AI
Aug 6

Reasoning Dynamics and the Limits of Monitoring Modality Reliance in Vision-Language Models

arXiv:2604. 14888v3 Announce Type: replace-cross Abstract: Recent advances in vision language models (VLMs) offer reasoning capabilities, yet how these unfold and integrate visual and textual information remains unclear.

By Danae S\'anchez Villegas, Samuel Lewis-Lim, Nikolaos Aletras, Desmond Elliott
llmsmultimodalsafety
More like this →
arXiv AI
Jul 16

Accuracy Without Grounding: Diagnosing Visual Dependency Dissociation in Video LLM Benchmarks

arXiv:2607. 13305v1 Announce Type: cross Abstract: Benchmark accuracy in video large language models (LLMs) is often treated as evidence of visual understanding.

By Jae Joong Lee
llmsbenchmarks
More like this →
arXiv Computation and Language
Aug 25

Knowing Isn't Always Saying: When Do Spatial Encodings Reach Answers in Vision-Language Models?

arXiv:2608.22916v1 Announce Type: new Abstract: Vision-language models are known to encode spatial information in their hidden states, yet often fail to use it when answering. However, it remains unc...

By Zeyu Wang, Xinming Xu
llmsmultimodal
More like this →
arXiv AI
Jun 17

Visuals Lie, Consistency Speaks: Disentangling Spatial Attention from Reliability in Vision-Language Models

arXiv:2606. 17389v1 Announce Type: cross Abstract: Multimodal Foundation Models are increasingly used as reasoning agents, making reliability, knowing when a model may hallucinate, critical.

By Logan Mann, Yi Xia, Ajit Saravanan, Ishan Dave, Saadullah Ismail, Shikhar Shiromani, Emily Huang, Ruizhe Li, Kevin Zhu
llmsagentsmultimodal
More like this →
arXiv AI
Jun 29

Dismantling Pathological Shortcuts: A Causal Framework for Faithful LVLM Decoding

arXiv:2606. 27596v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) exhibit sophisticated reasoning but remain susceptible to object hallucination.

By Liu Yu, Can Chen, Ping Kuang, Zhikun Feng, Fan Zhou, Gillian Dobbie
llmsmultimodalbenchmarkssafety
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea