arXiv:2607. 21806v1 Announce Type: new Abstract: Predictive machine learning (ML) models are increasingly used to aid human decision-makers across various high-risk domains such as healthcare and criminal justice.
By Jonathan Zhang, Erik Skalnes, Jacob Chen, Michael Oberst
arXiv:2605. 05882v2 Announce Type: replace-cross Abstract: Artificial-intelligence systems are becoming ubiquitous in society, yet their predictions typically inherit biases with respect to protected attributes such as race, gender, or age.
By Filip Edstr\"om, Guilherme W. F. Barros, Tetiana Gorbach, Xavier de Luna
arXiv:2606. 02198v1 Announce Type: new Abstract: Prediction tasks over individual futures, which are inherently noisy, often admit multiple similarly accurate models.
By Ashwin Singh, Carlos Castillo
arXiv:2607. 26908v1 Announce Type: new Abstract: In many domains such as Palliative Care, Credit Assignment and Recommender Systems, predictions may causally influence the outcomes they predict.
By Brandon Gower-Winter, Georg Krempl
In many domains such as Palliative Care, Credit Assignment and Recommender Systems, predictions may causally influence the outcomes they predict. This phenomena is known as Outcome Performativity.
arXiv:2505. 08908v3 Announce Type: replace-cross Abstract: Many researchers apply classical statistical decision theory to evaluate treatment choices and learn optimal policies.
By Benedikt Koch, Kosuke Imai
arXiv:2607. 20787v1 Announce Type: cross Abstract: For two decades, the standard remedy for class-imbalanced learning has been to fabricate synthetic minority examples, and the standard evidence of their validity has been a check that cannot fail: synthetic points are scored against the very data that generated them.
By Ahmad B. Hassanat, Ahmad S. Tarawneh, Ghada A. Altarawneh
arXiv:2606. 05029v1 Announce Type: new Abstract: Controlled experiments are the backbone of machine learning research, but at the scale of modern foundation models, they have become prohibitively expensive.
By Gunnar K\"onig, Martin Pawelczyk, Ulrike von Luxburg, Sebastian Bordt
arXiv:2606. 16723v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly take actions (screening applicants, recommending credit, triaging patients), yet fairness for LLMs is still measured by grading answers.
By Triveni Morla, Rohith Reddy Bellibaltu, Manpreet Singh, Manmeet Singh Kapoor
arXiv:2604. 23904v3 Announce Type: replace-cross Abstract: Synthetic tabular data are often evaluated by distributional similarity, privacy distance, or train-on-synthetic-test-on-real predictive performance, but these criteria do not ensure validity for causal inference.
By Yichen Xu
arXiv:2606. 08275v1 Announce Type: cross Abstract: When an LLM agent fails -- issues a refund it should not have, calls the wrong tool, leaks data -- existing tooling answers what happened (observability) or whether it passed (evaluation), but not which step caused the failure.
By Jaineet Shah
arXiv:2608. 02877v1 Announce Type: new Abstract: Causal discovery recovers directed structure from observational data and is increasingly used in clinical settings to support mechanism reasoning and fairness audits of predictive models.
By Nitish Nagesh, Elahe Khatibi, Thomas Dean Hughes, Mahdi Bagheri, Pratik Gajane, Amir M. Rahmani