arXiv AI

Off-Policy Evaluation with Strategic Agents via Local Disclosure

arXiv:2606. 07308v1 Announce Type: new Abstract: We study off-policy evaluation (OPE) under strategic behavior where decision subjects (or agents) respond to a decision maker's policy by strategically modifying their covariates.

Hugging Face Trending Papers
Jun 28

Safety from Honesty in a Disinterested AI Predictor

As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified. We present a formal safety argument for the Scientist AI (SAI) Predictor, trained to approximate the Bayesian posterior conditioned on a dataset of "epistemically contextualized" natural-language statements.

arXiv AI
Jun 30

Safety from Honesty in a Disinterested AI Predictor

arXiv:2606. 29657v1 Announce Type: new Abstract: As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified.

By Yoshua Bengio, Oliver Richardson, Tom\'a\v{s} Gaven\v{c}iak, Michael Cohen, Rory Svarc, Damiano Fornasiere, Gael Gendron, David Hyland, Aton Kamanda, Adam Oberman, Francis Rhys Ward, Anna Gaven\v{c}iak, Jacob Livingston Slosser, Vincent Mai, Iulian Serban, Joumana Ghosn
arXiv AI
6d ago

Which Influence Are We Estimating? The Role of Counterfactual Specifications in Data Attribution

The paper investigates how the definition of influence—specifically the behavior being attributed, the intervention on training data, and the counterfactual training process—affects rankings produced by influence estimators. It formalizes influence as a counterfactual estimand, distinguishes specification mismatch from approximation error, and categorizes existing estimators by their implied specifications. Experiments demonstrate that different specifications can lead to markedly different rankings, and that careful specification choice improves attribution quality in tasks such as noisy label detection and large‑language‑model attribution.

By Zhe Li, Wei Zhao, Peixin Zhang, Jun Sun
arXiv AI
Sep 25

When Honesty is Not Enough in AI Debate

The paper "When Honesty is Not Enough in AI Debate" explores how AI debate, intended as a scalable oversight method, can allow agents to pursue hidden objectives while still achieving correct verdicts. By introducing the strategic interactive oversight (SIO) framework, the authors formalise task‑admissible latent optimisation and demonstrate, via the establish protocol debate, a trade‑off between task success and disclosure of a hidden variable. They show that expanding the cross‑examiner’s role can reduce bias, underscoring that oversight effectiveness depends not only on verdict correctness but also on the information revealed in transcripts.

By Rayne Holland, Liming Zhu, Jason Xue
arXiv AI
Jun 9

Beyond Rational Illusion: Behaviorally Realistic Strategic Classification

arXiv:2605. 19674v2 Announce Type: replace Abstract: Strategic classification(SC) studies the interaction between decision models and agents who strategically manipulate their features for favorable outcomes.

By Xinpeng Lv, Yunxin Mao, Renzhe Xu, Chunyuan Zheng, Yikai Chen, Haoxuan Li, Yang Shi, Jinxuan Yang, Zhouchen Lin, Yuanlong Chen, Yuanxing Zhang, Shaowu Yang, Wenjing Yang, Haotian Wang
Hugging Face Trending Papers
Jun 24

RevengeBench: Reverse Engineering Code-Space Policies from Behavioral Experiments

For most of scientific history, researchers studying behavior could only infer hidden mechanisms from outward actions: an inverse problem that becomes more tractable when observation is augmented by targeted intervention. We pose a computational analogue: given only behavioral traces of an agent in a game environment, can a learner reconstruct the underlying decision program as executable code, and how much does this reconstruction improve with the ability to design controlled experiments?

arXiv Machine Learning
Jun 25

RevengeBench: Reverse Engineering Code-Space Policies from Behavioral Experiments

arXiv:2606. 26094v1 Announce Type: new Abstract: For most of scientific history, researchers studying behavior could only infer hidden mechanisms from outward actions: an inverse problem that becomes more tractable when observation is augmented by targeted intervention.

By Babak Rahmani, Sebastian Dziadzio, Joschka Str\"uber, Sergio Hern\'andez-Guti\'errez, Matthias Bethge
Hugging Face Trending Papers
Jun 29

Decision-Value Attribution in Predict-then-Optimize Systems

Predictive models are increasingly embedded in operational decision-making, yet standard explanation methods typically explain forecasts rather than the decisions those forecasts induce. This distinction is important in predict-then-optimize systems: large forecast changes may leave the optimizer's action unchanged, while small changes can alter the selected decision and its realized value.

arXiv Machine Learning
Jun 30

Decision-Value Attribution in Predict-then-Optimize Systems

arXiv:2606. 29878v1 Announce Type: new Abstract: Predictive models are increasingly embedded in operational decision-making, yet standard explanation methods typically explain forecasts rather than the decisions those forecasts induce.

By Konstantinos Ziliaskopoulos, Alexander Vinel, Alice E. Smith
arXiv AI
Aug 24

Why2Speak: Faithful Reasoning for Abstaining Action Policies

The paper investigates how agentic systems decide between acting and abstaining, focusing on the fidelity of their reasoning explanations. Using Qwen3‑8B in a multi‑party conversation setting, the authors compare direct decision policies, reasoning policies, supervised fine‑tuning, and reinforcement learning, finding a trade‑off: strong direct policies yield higher performance but no traceable reasoning, while reasoning policies provide an audit trail at the cost of lower recall. The study also uncovers that exposing reasoning can alter the agent’s policy and that common faithfulness metrics may overstate the alignment between reasoning and decisions.

By Shreya Mendi, Brinnae Bent