arXiv AI By Qiming Shi, Zhaolu Kang, Yunfan Zhou, Di Weng, Yingcai Wu

SPADER: Step-wise Peer Advantage with Diversity-Aware Exploration Rewards for Multi-Answer Question Answering

Read the original on arXiv AI →

arXiv:2606. 00593v1 Announce Type: cross Abstract: Large language models are increasingly deployed as tool-augmented agents to acquire information beyond parametric knowledge.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.