arXiv AI
Aug 24

Re$^3$Cap: Retrieval-Guided Refinement for Image Captioning Enhancement via Reinforcement Learning

Re$^3$Cap introduces a retrieval‑guided refinement strategy for image captioning that leverages multi‑modal retrieval as a reasoning signal. The method, built on a Caption Refinement Suggester and a Caption Quality Assessor, detects hallucinations and omissions to produce more accurate and detailed captions without extra annotations. Experiments show it surpasses supervised fine‑tuning and improves relation reasoning by 8.64% on the COCO‑LN500 benchmark.

By Haonan Jia, Shichao Dong, Zenghui Sun, Jiawen Zheng, Ziqi Miao, Gege Shi, Qiuyu Zhao, Jinsong Lan, Xiaoyong Zhu, Bo Zheng
arXiv AI
Jun 17

See First, Answer Later: Visual Evidence Pre-Alignment via Sufficiency-Driven RL

arXiv:2606. 17678v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) integrate strong text reasoning with visual inputs, yet their responses can be inconsistent with the underlying images, indicating ineffective utilization of visual evidence during inference.

By Yilian Liu, Sicong Leng, Guoshun Nan, Junyi Zhu, Jiayu Huang, Minghao Sun, Xuancheng Zhu, Yisong Chen, Zexian Wei, Xiaofeng Tao