arXiv Computer Vision By Peipei Li, Shuhan Xia, Shengyang Liu, Zekun Li, Ran He

MM-VeriAgent: Learning to Use Extensive Tools to Verify Multimodal Misinformation with Reinforcement Learning

Read the original on arXiv Computer Vision →

MM-VeriAgent is a reinforcement‑learning framework that learns to verify multimodal misinformation by leveraging a specialized toolkit called MM-VeriTools. The toolkit encapsulates the strongest models for textual, visual, and cross‑modal forgery analysis as callable tools with a unified interface. To improve training efficiency, the authors introduce a Tool‑Execution Cache that pre‑executes candidate tool calls and reuses cached outputs, resulting in substantial accuracy gains on MMFakeBench and reduced online tool executions during training.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computation and Language
Sep 16

Spurious Tool Use: When RL Agents Learn the Wrong Reason to Act

The paper investigates how reinforcement learning can cause large language model agents to adopt shortcut policies for tool use, relying on superficial prompt cues rather than actual task needs. By creating synthetic environments that mix factual QA and math reasoning, the authors show that agents often invoke tools when cues are present, even when those tools are unnecessary, with spurious invocation rates rising up to 39%. They find that shortcut learning occurs mainly when agents have already mastered the target tool and that semantic alignment between cues and tools amplifies the effect. To counter this, they propose a dense, decision-level reward where an LLM judge assesses tool necessity, which reduces cue-driven tool use while maintaining performance.

By Yiwei Yang, Haoxiang Zhang, Bingbing Wen, Yao Lu, Yuchen Wu, Lei Zhang, Julian McAuley, Pan Lu, Bill Howe