arXiv:2606. 30136v1 Announce Type: new Abstract: Humans facing algorithmic decision systems have been found to ``game'' them by altering their input data (at a cost to them) in order to favorably change the algorithmic outcomes they receive (at a cost to the algorithm).
By Sura Alhanouti, G\"uzin Bayraksan, Parinaz Naghizadeh
arXiv:2606. 28204v1 Announce Type: cross Abstract: Algorithmic developments in Strategic Classification have been mostly limited to linear classifiers in settings where the best response has a closed-form solution or can be easily approximated.
By Jack Geary, Boyan Gao, Henry Gouk
arXiv:2602. 11439v3 Announce Type: replace Abstract: Strategic classification studies the problem where self-interested individuals or agents manipulate their response to obtain favorable decision outcomes made by classifiers, typically turning to dishonest actions when they are less costly than genuine efforts.
By Ziyuan Huang, Lina Alkarmi, Mingyan Liu
arXiv:2609.39818v1 Announce Type: new
Abstract: In many high-stakes settings, human decision-makers can acquire support information before making a decision. However, acquiring information is costly,...
By Carlotta Giacchetta, Alessando Bogani, Cesare Barbera, Giovanni De Toni, Michele Caprio, Andrea Pugnana, Andrea Passerini
arXiv:2605. 19674v2 Announce Type: replace Abstract: Strategic classification(SC) studies the interaction between decision models and agents who strategically manipulate their features for favorable outcomes.
By Xinpeng Lv, Yunxin Mao, Renzhe Xu, Chunyuan Zheng, Yikai Chen, Haoxuan Li, Yang Shi, Jinxuan Yang, Zhouchen Lin, Yuanlong Chen, Yuanxing Zhang, Shaowu Yang, Wenjing Yang, Haotian Wang
arXiv:2606. 01198v1 Announce Type: new Abstract: Strategic classification studies settings in which agents respond to a deployed classifier by modifying observable features at a cost.
By Siddharth Shrivastava, Mahvith Akshintala, B Vamsha Vardhan Reddy, Naresh Manwani, Sujit Gujar, Ganesh Ghalme
arXiv:2606. 10347v1 Announce Type: new Abstract: Machine learning is increasingly used in critical domains, where both predictions and their associated confidence levels influence important decisions.
By Vin\'icius Peixoto Chagas, Carlos Henrique Leit\~ao Cavalcante, Thiago Alves Rocha
The paper introduces the concept of observational multiplicity, where multiple probabilistic classifiers can perform similarly yet produce conflicting predictions, undermining interpretability and safety. It proposes measuring this arbitrariness through a regret metric that captures how predictions could shift with different training labels. The authors present a general method to estimate regret, show it varies across dataset groups, and discuss its use for safety via abstention and targeted data collection.
By Erin George, Deanna Needell, Berk Ustun
The paper surveys 25 studies that use explainable AI to compromise machine learning models, covering attacks such as model extraction, membership inference, and model inversion. It distinguishes between how explanations are obtained—through target releases, attacker-derived methods, secondary disclosure, privileged access, or global artifacts—and shows that explanations can lower extraction costs and reveal membership signals via statistics, recourse distance, and robustness. The authors compare threat models, signals, and defenses, concluding that no single explanation type is always unsafe and that protection must be tailored to the specific acquisition path and target asset.
By Abdullah Caglar Oksuz, Anisa Halimi, Erman Ayday
As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified. We present a formal safety argument for the Scientist AI (SAI) Predictor, trained to approximate the Bayesian posterior conditioned on a dataset of "epistemically contextualized" natural-language statements.
arXiv:2606. 29657v1 Announce Type: new Abstract: As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified.
By Yoshua Bengio, Oliver Richardson, Tom\'a\v{s} Gaven\v{c}iak, Michael Cohen, Rory Svarc, Damiano Fornasiere, Gael Gendron, David Hyland, Aton Kamanda, Adam Oberman, Francis Rhys Ward, Anna Gaven\v{c}iak, Jacob Livingston Slosser, Vincent Mai, Iulian Serban, Joumana Ghosn
arXiv:2606. 00826v1 Announce Type: new Abstract: Strategic machine learning investigates scenarios where agents manipulate their features to receive favorable decisions from predictive models.
By Xinpeng Lv, Chunyuan Zheng, Yunxin Mao, Renzhe Xu, Hao Zou, Shanzhi Gu, Liyang Xu, Huan Chen, Yuanlong Chen, Wenjing Yang, Haotian Wang