arXiv AI By Sam Ryan

Durable Evaluation Framework: Adversarial Arbitration for Sycophancy Reduction in Large Language Models

Read the original on arXiv AI →

arXiv:2606. 07532v2 Announce Type: replace-cross Abstract: RLHF-trained models are systematically biased toward agreement over accuracy, a structural property of the training process.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

arXiv AI
Aug 11

Unified Hallucination Fuzzing for Multimodal Large Language Models

arXiv:2608. 07525v1 Announce Type: cross Abstract: Hallucination remains a persistent challenge for Multimodal Large Language Models (MLLMs), severely limiting their reliability in high-stakes applications.

By Pengfei Zhou, Jiajun Song, Zhiwei Tang, Yixing Ma, Xiaopeng Peng, Donghui Si, Yuhang Xu, Huiqi Song, Yiyuan Miao, Yichen Qian, Weihua Chen, Wangbo Zhao, Bohan Zhuang, Jiasheng Tang, Yang You