arXiv:2606. 26523v1 Announce Type: new Abstract: We develop a framework for interpreting AI systems as agents, drawing on the philosophical tradition of radical interpretation and the tools of mechanistic interpretability.
By Daniel A. Herrmann, Benjamin A. Levinstein
arXiv:2606. 12289v1 Announce Type: cross Abstract: As Artificial Intelligence models grow in complexity, interpretability has become an indispensable tool for understanding, debugging, and controlling their computations.
By Pietro Barbiero, Giovanni De Felice, Mateo Espinosa Zarlenga, Francesco Giannini, Filippo Bonchi, Mateja Jamnik, Giuseppe Marra, Ruggero Noris
arXiv:2606. 26228v1 Announce Type: cross Abstract: We review the concepts of interpretability and explainability as they apply to machine learning in physics.
By Rikab Gambhir, Luisa Lucie-Smith, Jesse Thaler
arXiv:2606. 11769v1 Announce Type: new Abstract: The European AI Act is the first comprehensive regulation of artificial intelligence (AI), setting out extensive obligations, particularly for so-called high-risk and general-purpose AI systems.
By Maximilian Poretschkin, Tabea Naeven
arXiv:2605. 27618v2 Announce Type: replace Abstract: Despite the wide use of explainability techniques to attempt to understand the behavior of Artificial Intelligence (AI), the generated explanations may not always be reliable.
By Tom\'as Pereira, Jo\~ao Vitorino, Eva Maia, Isabel Pra\c{c}a
arXiv:2608. 02238v1 Announce Type: cross Abstract: Ensuring trust in AI systems is essential for the safe and ethical integration of machine learning systems into high-stakes domains such as digital health.
By Abdullah Mamun, Shovito Barua Soumma, Hassan Ghasemzadeh
arXiv:2607. 07229v1 Announce Type: new Abstract: Prior work has shown that chain-of-thought (CoT) reasoning is often unfaithful: a model's stated reasoning does not reliably reflect the process that produced its output.
By Silvia Santano
arXiv:2607. 21209v1 Announce Type: cross Abstract: In the field of Artificial Intelligence, an agent is a system which is able to autonomously make decisions in order to reach a desired goal.
By Heather Merhout (Miami University), Daniela Inclezan (Miami University)
arXiv:2410. 22526v2 Announce Type: replace Abstract: To effectively address potential harms from Artificial Intelligence (AI) systems, it is essential to identify and mitigate system-level hazards.
By Shalaleh Rismani, Roel Dobbe, AJung Moon
arXiv:2608. 12325v1 Announce Type: new Abstract: Autonomous reasoning is among the most scientifically and economically motivating topics in AI today.
By Rachel Lawrence, Jacqueline Maasch
arXiv:2607. 15394v1 Announce Type: new Abstract: Black-box models limit the adoption of artificial intelligence in medicine due to their lack of interpretability and reproducibility.
By Antony Garcia, Adrian Noriega, Gabrielle Britton, Xinming Huang
arXiv:2606. 16786v1 Announce Type: new Abstract: Algorithmic explanations are intended to help stakeholders understand opaque algorithmic decisions, but in practice, they often fall short.
By Eric G\"unther, Bal\'azs Szabados, Kristof Meding, Gunnar K\"onig, Sebastian Bordt, Ulrike von Luxburg