arXiv:2607. 00164v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards can in principle train calibrated probabilistic forecasters, since a proper scoring rule such as the Brier score is computed from outcomes alone and is minimized in expectation by the true probability.
By Sadanand Singh, Allam Reddy, Manan Chopra
arXiv:2607. 13618v1 Announce Type: new Abstract: LLM agents are increasingly evaluated on multi-week decision tasks in which the state that drives cost is never directly observed.
By Sagar Deb, Ashwanth Krishnan
LLM agents are increasingly evaluated on multi-week decision tasks in which the state that drives cost is never directly observed. On such tasks the final cost cannot say why an agent failed: it may have misread the world, or read it correctly and still failed to act (the knowing-doing gap).
arXiv:2607. 04419v2 Announce Type: replace Abstract: Final-answer scores hide which agent transitions helped or harmed a trace.
By Andrew Zhang, Chengzhan Li
arXiv:2603. 02491v3 Announce Type: replace-cross Abstract: As artificial agents become increasingly capable, what internal structure is necessary for an agent to act competently under uncertainty?
By Aran Nayebi
arXiv:2606. 12200v1 Announce Type: cross Abstract: We study policy representation learning from unlabeled multi-policy behavioral data.
By Andrew Kang, Priya Narasimhan
arXiv:2608. 01587v1 Announce Type: cross Abstract: Machine-learning benchmarks often pair a label that aggregates a long temporal horizon with input observed through one or a few short windows.
By Xizhe Zhang
arXiv:2607. 11607v1 Announce Type: new Abstract: Distributional reinforcement learning agents learn full return distributions that are increasingly read at face value: for interpretability, risk-sensitive control, and safety monitoring.
By Hari Prasad
arXiv:2607. 05904v1 Announce Type: new Abstract: Training a language model against its own reference-free judgments (the premise of self-rewarding, self-play, and LLM-as-a-judge pipelines) assumes a model's verdict on a shown answer tracks correctness.
By Chenyu Zhou
arXiv:2607. 10814v1 Announce Type: cross Abstract: Evaluating LLM agents in hidden-information multi-agent settings is hard: final outcomes are high-variance and rarely reveal why an agent decided as it did.
By Yuan Gao, Jiangyi Yang, Yao Zhao, Yichi Zhang
arXiv:2608. 14982v1 Announce Type: cross Abstract: Transformers applied to spatial imperfect-information games must represent map geometry while tracking hidden entities through time.
By Wenji Fu
arXiv:2607. 10362v1 Announce Type: new Abstract: Latent world models are trained to predict future states in a learned representation and are then deployed inside a planner that selects actions by simulating them forward.
By Hanzhe You, Yonggang Zhang, Maohao Ran, Zhiqin Yang, Zhenyuan Zhang, Wei Xue, Jun Song, Xinmei Tian, Yike Guo