arXiv AI

MMPCBench: Benchmarking Multimodal Large Language Models on Proactive Critique of Flawed Inputs

arXiv Computation and Language
Aug 27

From Passive Response to Proactive Correction: Enhancing LLM Robustness Against Input Fact Perturbations

The paper introduces DEDUCE, a three‑stage framework that turns large language models into proactive error correctors by detecting input fact errors, devising correction strategies, and delivering reliable answers. It also presents MisFactQA, a dataset of factual errors, and new metrics for robustness evaluation. Experiments on TruthfulQA, FalseQA, and MisFactQA show significant gains in accuracy and error correction across Qwen, LLaMA, and Gemma models.

By Ping Wang, Xiangguo Sun, Bingbing Xu, Guocong Li, Xiaofeng Meng
arXiv Computer Vision
Aug 26

DoublesEval: Diagnosing Multi-Agent Tactical Reasoning in Vision-Language Models via Professional Doubles Badminton

The paper introduces DoublesEval, a diagnostic framework that uses professional doubles badminton to test visual‑language models’ ability to reason about dynamic multi‑agent interactions. It decomposes rallies into key moments and evaluates models across four dimensions—atomic recognition, intra‑segment composite understanding, cross‑segment causal reasoning, and high‑level tactical abstraction—highlighting specific reasoning failures. The authors also propose TacticCheck, a lightweight consistency checker that improves performance without retraining the models, yet significant gaps remain in tactical reasoning.

By Jintao Cheng, Weibin Li
arXiv AI
Aug 17

The Metacognitive Bottleneck: Japanese Riddles Reveal Fundamental Limits of Machine Insight and Self-Evaluation in Reasoning AI

arXiv:2509. 14704v3 Announce Type: replace Abstract: Benchmark saturation and training-data contamination increasingly obscure whether reported gains in large language models (LLMs) reflect genuine advances in reasoning or familiarity with recurring patterns in benchmark problems.

By Masaharu Mizumoto, Dat Nguyen, Zhiheng Han, Xingfu Li, Yo Nakawake, Le Minh Nguyen
arXiv AI
3d ago

Evidence-Grounded Trustworthy Multimodal Reasoning and Evaluation Benchmark in Complex Urban Scenes

arXiv:2506.09557v2 Announce Type: replace-cross Abstract: While Multimodal Large Language Models (MLLMs) demonstrate impressive performance in benign scenarios, their cognitive reliability deteriorat...

By Zhaoyang Wei, Bowen Jiang, Xumeng Han, Jiashu Li, Xuehui Yu, Yuling Liu, Guorong Li, Zhenjun Han, Jianbin Jiao
arXiv AI
1d ago

Multimodal Language Models as Text-to-Image Model Evaluators

Multimodal Language Models as Text-to-Image Model Evaluators presents MT2IE, a framework where a multimodal large language model generates evaluation prompts and scores images, achieving higher correlation with human judgment than prior metrics. MT2IE recovers official T2I model rankings using only 20 prompts—far fewer than traditional benchmarks—and adapts prompts to each model’s performance, maintaining informative scoring ranges. The approach demonstrates that dynamic, interactive evaluation can replace static benchmarks as T2I models improve.

By Jiahui Chen, Candace Ross, Reyhane Askari-Hemmat, Koustuv Sinha, Melissa Hall, Amy Zhang, Michal Drozdzal, Adriana Romero-Soriano