arXiv AI

The Benchmark Trap: Structures of Power and Injustice in AI Evaluations

arXiv:2608. 15326v1 Announce Type: new Abstract: Artificial intelligence (AI) benchmarks are not neutral tools of evaluation but socio-technical artefacts that shape competition, power, and research priorities within AI.

arXiv AI
Jun 16

Artificial Intelligence Index Report 2026

arXiv:2606. 15708v1 Announce Type: new Abstract: Welcome to the ninth edition of the AI Index report.

By Sha Sajadieh, Loredana Fattorini, Raymond Perrault, Yolanda Gil, Vanessa Parli, Lapo Santarlasci, Juan Pava, Nestor Maslej, Russ Altman, Erik Brynjolfsson, Carla Brodley, Jack Clark, Virginia Dignum, Vipin Kumar, James Landay, Terah Lyons, James Manyika, Juan Carlos Niebles, Yoav Shoham, Elham Tabassi, Russell Wald, Toby Walsh, Dan Weld
arXiv Machine Learning
Sep 11

From Protocols to Evidence: Bounded Claims for AI in Service of the Common Good

The paper argues that AI should be evaluated not only by principles but by concrete protocols that translate commitments into roles, requirements, records, oversight, and assessment. It introduces a rupture test linking institutional baselines to system evaluation, and distinguishes evidence‑bounded deployment from measurement‑bounded governance. The authors propose the RISE AI architecture to make bounded, evidence‑based claims about Responsibility, Inclusivity, Safety, and Empowerment, emphasizing the need for engineering, institutional repair, and ongoing moral judgment.

By Nitesh V. Chawla, Paulo Benanti
arXiv AI
Aug 11

Abstracted Away: Resisting Alienation and Ungrounded Abstraction in AI Research Communities

arXiv:2608. 08408v1 Announce Type: cross Abstract: Logics of abstraction in computational AI research often push important forms of knowledge and reflection aside: dominant standards of legitimacy separate from lived experience of harm; the goals of work misalign with the practices that operationalize them; and career demands crowd out critical reflection.

By Vyoma Raman, Isabel O. Gallegos, Neha Srivathsa
arXiv AI
Jul 1

A Technical Typology of AI Systems in Public Administration

arXiv:2606. 31755v1 Announce Type: cross Abstract: Research on artificial intelligence (AI) in the public sector often treats "AI" as a single category, neglecting technical distinctions between different AI systems.

By Jonathan Rystr{\o}m, Chris Schmitz, Nathan Davies, Gerhard Hammerschmid, Albert Meijer, Chris Russell
arXiv AI
Jun 9

Can Data Work be Reparative?

arXiv:2606. 09408v1 Announce Type: cross Abstract: We present an ethnographic study of an alternative approach to data work, developed by a civic-tech initiative that builds datasets for training and benchmarking online safety systems.

By Srravya Chandhiramowuli, Ding Wang, Alex Taylor
arXiv AI
Aug 18

Position: AI Lock-In Is in Progress, and We Must Be Prepared

arXiv:2608. 14565v1 Announce Type: new Abstract: AI safety research has mainly focused on two areas: technical alignment (ensuring AI systems produce human-aligned outputs) and the regulation of generative AI's societal impacts (including unemployment risk and labor market disruption).

By Jaeho Kim, Seokhyun Lee, Jieun Lee, Changhee Lee
arXiv AI
Jul 17

Global Index on Responsible AI: 2026 Report

arXiv:2607. 14782v1 Announce Type: new Abstract: Grounded in human rights-based frameworks such as the UNESCO Recommendation on the Ethics of AI, the Global Index on Responsible AI (GIRAI) examines how countries translate responsible AI commitments into enforceable protections, institutional capacity, and redress mechanisms.

By Rachel Adams, Fola Adeleke, Ayantola Alayande, Selamawit Engida Abdella, Ana Florido, Nicol\'as Grossman, Leah Junck
arXiv AI
Sep 23

Biased AI improves human performance but reduces perceived helpfulness

The study tests deliberately biased AI assistants and finds that such bias improves human performance on tasks like misinformation evaluation, financial investment, and graduate education compared to neutral AI. However, participants undervalue biased AI and overvalue neutral AI, even when performance is similar. When two AI biases flank a participant’s perspective, performance gains are maintained while reducing the perceived cost and one‑sided influence.

By Shiyang Lai, Jiwoong Choi, Junsol Kim, Nadav Kunievsky, Yujin Potter, James Evans