arXiv:2609.17296v1 Announce Type: cross
Abstract: Policy learning aims to determine who should be treated based on individual characteristics. In high-stakes settings such as medicine and public poli...
By Ying Jin, Naoki Egami
arXiv:2607. 02206v1 Announce Type: cross Abstract: Predictions are increasingly used to guide high-stakes decisions, from treatment selection to policy making.
By Yurui Zheng, Ying Jin
arXiv:2606. 13884v1 Announce Type: new Abstract: Modern decision systems increasingly rely on learned components whose outputs may be confident yet wrong, exposing downstream actions to costly errors.
By Laxmipriya Ganesh Iyer, Rahul Suresh Babu
arXiv:2607. 05620v1 Announce Type: cross Abstract: In many decision-making settings, new interventions are acceptable only if they do not reduce outcomes below some established threshold.
By Katherine Avery, Bruno Castro da Silva, David Jensen
arXiv:2603. 02196v3 Announce Type: replace Abstract: An agent must try new behaviors to explore and improve.
By Drew Prinster, Clara Fannjiang, Ji Won Park, Kyunghyun Cho, Anqi Liu, Suchi Saria, Samuel Stanton
arXiv:2606. 07399v1 Announce Type: cross Abstract: Generative models for counterfactual outcomes have great potential to support decision-making under complex interventions, but existing approaches are limited by unstable estimation, poor generalization across environments, and bias from nuisance model misspecification.
By Raphael C Kim, Jingsen Zhu, Ramin Zabih, Michele Santacatterina
arXiv:2505. 08908v3 Announce Type: replace-cross Abstract: Many researchers apply classical statistical decision theory to evaluate treatment choices and learn optimal policies.
By Benedikt Koch, Kosuke Imai
arXiv:2609.15254v1 Announce Type: cross
Abstract: Conformal counterfactual prediction constructs prediction sets with finite-sample coverage guarantees for counterfactual outcomes and individual trea...
By Matteo Zecchin, Osvaldo Simeone
arXiv:2608. 08743v1 Announce Type: cross Abstract: Reinforcement learning (RL) seeks to optimize sequential decisions to maximize population-level benefits over time.
By Jianhan Zhang, Jitao Wang, John D. Piette, Donglin Zeng, Chengchun Shi, Zhenke Wu
arXiv:2606. 05614v1 Announce Type: new Abstract: Large language models (LLMs) are rigorously aligned to refuse harmful requests, a process that inherently cultivates a latent capacity to evaluate and recognize unsafe content.
By Long P. Hoang, Hai V. Le, Shaoyang Xu, Wei Lu, Wenxuan Zhang
arXiv:2606. 06391v1 Announce Type: cross Abstract: Sharing the financial impact of rare adverse events across a group can soften extreme individual burdens, but any participant made worse off by the arrangement has reason to leave.
By Ieva Kazlauskaite
RePolicy is a reinforcement learning approach designed to invoke safety policies for language model agents by evaluating entire execution trajectories within context-dependent policy libraries. It generates policy-grounded rationales and safety judgments, and is initialized with the PolicyTraj-20K dataset before fine-tuning via GRPO with verifiable rewards and policy-context perturbation. Experiments on six safety benchmarks demonstrate strong safety-detection performance and robust policy invocation across varying contexts.
By Houcheng Jiang, Boxuan Zhang, Qiyong Zhong, Junfeng Fang, Xiang Wang, Xiangnan He