← Back to all news
arXiv AI September 23, 2026 By Kunyu Peng, Junming Liu, Ruiqi He, Qingzhuo Wang, Jianzhong Qi, Xianhui Liu

When Should a VLM Look? Paying Only for Visual Calls That Were Needed and Used

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

  • agents
  • multimodal
  • benchmarks

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
2d ago

Visual sensitivity is not claim retractability: persistence-aware credit assignment for multimodal reinforcement learning

arXiv:2609.36572v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has been extended to Large Vision-Language Models (LVLMs), and perception-aware methods further e...

By Zhongan Bi, Kepeng Lin, Xuanang Gao, Yuhan Sun, Lianrun Zhang
llmsreinforcement-learningmultimodalbenchmarks
More like this →
arXiv Machine Learning
Sep 24

ProCredit: From Outcome Rewards to Progress Credit in Agentic Reinforcement Learning

arXiv:2609.27532v1 Announce Type: new Abstract: Long-horizon agentic tasks require an agent to modify an environment through a sequence of tool calls, with success determined by the final state. The...

By Ming Ma, Yi Zhu, Yiran Zhong, Feida Zhu, Chonghan Liu, Pengkun Jiao, Qichao Wang, Yanhao Jia, Tianming Yang, Steven Hoi
agentsreinforcement-learning
More like this →
arXiv AI
Jun 16

Looking Is Not Picking: An Attention-Segment Account of Tool-Selection Failures in LLM Agents

arXiv:2606. 16364v1 Announce Type: new Abstract: LLM agents mis-call tools, and the natural guess is that the model failed to see the right tool in a crowded harness.

By Shiyang Chen
llmsagentssafety
More like this →
arXiv Machine Learning
Sep 15

GroundBench: A Factorized, Counterfactual Benchmark for Locating VLM Affordance Failures

arXiv:2609.13308v1 Announce Type: cross Abstract: A companion evaluation found that naming the target part in a manipulation prompt increased action accuracy by 0.32-0.63 across eight vision-language...

By Sarthak Sattigeri
llmsroboticsmultimodalbenchmarks
More like this →
arXiv AI
Jun 9

Decision-Aware Memory Cards: Counterfactual-Inspired Context Selection and Compression for Tool-Using LLM Agents

arXiv:2606. 08151v1 Announce Type: new Abstract: Tool-using LLM agents often fail not because relevant text is absent, but because decisive evidence is not selected, compressed, or surfaced at action time.

By Xinyu Guan, Qianyang Zhao, Yuming Deng
llmsragagentsfine-tuningbenchmarks
More like this →
arXiv AI
1d ago

Evaluating Agents Across Runtime Contracts: When Mismatch Costs Efficiency or Quality

arXiv:2603.01209v3 Announce Type: replace Abstract: In CodeAct, language-model agents write Python that calls tools and use execution feedback to choose actions. Persistent runtimes preserve Python v...

By Victor May, Van Khue Nguyen, Aaditya Salgarkar, Yishan Wang, Diganta Misra, Huu Nguyen
agentsfine-tuning
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea