arXiv Machine Learning

Ambiguous Strategic Classification

arXiv:2606. 10137v1 Announce Type: new Abstract: A common assumption in strategic classification is that the classifier is public knowledge.

arXiv Machine Learning
Jun 29

Non-Linear Strategic Classification Made Practical

arXiv:2606. 28204v1 Announce Type: cross Abstract: Algorithmic developments in Strategic Classification have been mostly limited to linear classifiers in settings where the best response has a closed-form solution or can be easily approximated.

By Jack Geary, Boyan Gao, Henry Gouk
arXiv Machine Learning
Aug 4

Multi-Level Strategic Classification: Incentivizing Improvement through Promotion and Relegation Dynamics

arXiv:2602. 11439v3 Announce Type: replace Abstract: Strategic classification studies the problem where self-interested individuals or agents manipulate their response to obtain favorable decision outcomes made by classifiers, typically turning to dishonest actions when they are less costly than genuine efforts.

By Ziyuan Huang, Lina Alkarmi, Mingyan Liu
arXiv AI
Jun 9

Beyond Rational Illusion: Behaviorally Realistic Strategic Classification

arXiv:2605. 19674v2 Announce Type: replace Abstract: Strategic classification(SC) studies the interaction between decision models and agents who strategically manipulate their features for favorable outcomes.

By Xinpeng Lv, Yunxin Mao, Renzhe Xu, Chunyuan Zheng, Yikai Chen, Haoxuan Li, Yang Shi, Jinxuan Yang, Zhouchen Lin, Yuanlong Chen, Yuanxing Zhang, Shaowu Yang, Wenjing Yang, Haotian Wang
arXiv Machine Learning
Sep 16

Observational Multiplicity

The paper introduces the concept of observational multiplicity, where multiple probabilistic classifiers can perform similarly yet produce conflicting predictions, undermining interpretability and safety. It proposes measuring this arbitrariness through a regret metric that captures how predictions could shift with different training labels. The authors present a general method to estimate regret, show it varies across dataset groups, and discuss its use for safety via abstention and targeted data collection.

By Erin George, Deanna Needell, Berk Ustun
arXiv Machine Learning
Sep 11

SoK: Privacy Attacks on Machine Learning via Explainable AI

The paper surveys 25 studies that use explainable AI to compromise machine learning models, covering attacks such as model extraction, membership inference, and model inversion. It distinguishes between how explanations are obtained—through target releases, attacker-derived methods, secondary disclosure, privileged access, or global artifacts—and shows that explanations can lower extraction costs and reveal membership signals via statistics, recourse distance, and robustness. The authors compare threat models, signals, and defenses, concluding that no single explanation type is always unsafe and that protection must be tailored to the specific acquisition path and target asset.

By Abdullah Caglar Oksuz, Anisa Halimi, Erman Ayday
Hugging Face Trending Papers
Jun 28

Safety from Honesty in a Disinterested AI Predictor

As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified. We present a formal safety argument for the Scientist AI (SAI) Predictor, trained to approximate the Bayesian posterior conditioned on a dataset of "epistemically contextualized" natural-language statements.

arXiv AI
Jun 30

Safety from Honesty in a Disinterested AI Predictor

arXiv:2606. 29657v1 Announce Type: new Abstract: As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified.

By Yoshua Bengio, Oliver Richardson, Tom\'a\v{s} Gaven\v{c}iak, Michael Cohen, Rory Svarc, Damiano Fornasiere, Gael Gendron, David Hyland, Aton Kamanda, Adam Oberman, Francis Rhys Ward, Anna Gaven\v{c}iak, Jacob Livingston Slosser, Vincent Mai, Iulian Serban, Joumana Ghosn