arXiv AI By Muyu He, Anand Kumar, Tsach Mackey, Meghana Rajeev, James Zou, Nazneen Rajani

Impatient Users Confuse AI Agents: High-fidelity Simulations of Human Traits for Testing Agents

Read the original on arXiv AI →

arXiv:2510. 04491v3 Announce Type: replace Abstract: Despite rapid progress in building conversational AI agents, robustness is still largely untested.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

arXiv AI
Aug 11

Unified Hallucination Fuzzing for Multimodal Large Language Models

arXiv:2608. 07525v1 Announce Type: cross Abstract: Hallucination remains a persistent challenge for Multimodal Large Language Models (MLLMs), severely limiting their reliability in high-stakes applications.

By Pengfei Zhou, Jiajun Song, Zhiwei Tang, Yixing Ma, Xiaopeng Peng, Donghui Si, Yuhang Xu, Huiqi Song, Yiyuan Miao, Yichen Qian, Weihua Chen, Wangbo Zhao, Bohan Zhuang, Jiasheng Tang, Yang You
arXiv AI
Aug 5

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation

arXiv:2608. 03166v1 Announce Type: new Abstract: Role-Playing Language Agents (RPLAs) are increasingly deployed in high-stakes applications such as healthcare assistance, customer support, and education, where maintaining consistent personas, ethical constraints, and behavioral coherence under adversarial pressure is critical.

By Saqib Shouqi, Abdullah Nazly, Januki Wanniarachchi, Ravisha De Alwis