arXiv:2607. 00164v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards can in principle train calibrated probabilistic forecasters, since a proper scoring rule such as the Brier score is computed from outcomes alone and is minimized in expectation by the true probability.
By Sadanand Singh, Allam Reddy, Manan Chopra
arXiv:2607. 13618v1 Announce Type: new Abstract: LLM agents are increasingly evaluated on multi-week decision tasks in which the state that drives cost is never directly observed.
By Sagar Deb, Ashwanth Krishnan
LLM agents are increasingly evaluated on multi-week decision tasks in which the state that drives cost is never directly observed. On such tasks the final cost cannot say why an agent failed: it may have misread the world, or read it correctly and still failed to act (the knowing-doing gap).
arXiv:2607. 04419v2 Announce Type: replace Abstract: Final-answer scores hide which agent transitions helped or harmed a trace.
By Andrew Zhang, Chengzhan Li
arXiv:2603. 02491v3 Announce Type: replace-cross Abstract: As artificial agents become increasingly capable, what internal structure is necessary for an agent to act competently under uncertainty?
By Aran Nayebi
arXiv:2606. 12200v1 Announce Type: cross Abstract: We study policy representation learning from unlabeled multi-policy behavioral data.
By Andrew Kang, Priya Narasimhan