arXiv AI By Bart Jaworski

The AI Risk Observatory: What Can We Learn from AI Disclosures in Annual Reports About Societal Resilience?

Read the original on arXiv AI →

The study examines whether annual reports can reveal how companies disclose AI-related risks and responses. Using a two-stage classification pipeline on 9,821 reports from 1,362 UK listed firms (2020‑2026), the authors find that mentions of AI risk rose from 2.8% to 41.2% and AI adoption disclosures from 13.8% to 45.2%, with most risk mentions clustering around major vendors like Microsoft. Disclosure varies by sector and market segment, with Critical National Infrastructure and AIM reports lagging, and substantive risk disclosures remain rare—only 4.3% in 2025. "Why It Matters": The findings show that while AI risk is increasingly referenced in corporate reports, substantive disclosures are scarce, highlighting a gap in transparency that could affect societal resilience.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 9

RiskNet: A large-scale dataset of AI risk incidents from news with alignment and multi-dimensional annotations

arXiv:2606. 08376v1 Announce Type: cross Abstract: As artificial intelligence (AI) systems are increasingly deployed across socially consequential domains, reports of AI-related harms and failures have grown in frequency and diversity.

By Leihan Zhang, Wecheng Ye, Xianlong Ma, Haochuan Liu, Yang Li, Qianyu Zhang, Jinliang Chen, Qiang Yan
arXiv AI
Jun 9

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting

arXiv:2606. 09809v1 Announce Type: new Abstract: AI evaluation results are produced at scale but reported inconsistently across leaderboards, model cards, benchmark papers, and company blogs.

By Avijit Ghosh, Anka Reuel, Jenny Chim, Wm. Matthew Kennedy, Srishti Yadav, Jennifer Mickel, Yanan Long, Andrew Tran, Anastassia Kornilova, Damian Stachura, Kevin Klyman, Felix Friedrich, Jeba Sania, Max Lamparth, Jan Batzner, Anoop Mishra, Eliya Habba, Yixiong Hao, Nathan Heath, Shalaleh Rismani, Usman Gohar, Andrea Loehr, David Manheim, Ruchira Dhar, Sree Harsha Nelaturu, Aarush Sinha, Leshem Choshen, Drishti Sharma, Ishan Khire, Amit Saha, Subramanyam Sahoo, Michael Hardy, Michael Alexander Riegler, Kabir Manghnani, Michelle Lin, Yanan Jiang, Yilin Huang, Asaf Yehudai, Jessica Ji, Aris Hofmann, Mubashara Akhtar, Nuno Moniz, Yacine Jernite, Stella Biderman, Zeerak Talat, Sanmi Koyejo, Mykel Kochenderfer, Irene Solaiman
arXiv Computation and Language
Aug 25

Expectations and Practices around AI Disclosure in CS Research

The paper examines AI disclosure policies in top computer science venues, finding them to be highly under‑specified. A survey of 109 researchers shows that disclosures are deemed most necessary for research design tasks and when human involvement is low, and it compiles researchers’ expectations for disclosure content. Analysis of 13,867 disclosure statements from EMNLP 2025 and ICLR 2026 reveals a significant mismatch between these expectations and actual practice, such as frequent disclosure of writing assistance despite it being considered less necessary.

By Arati Mohapatra, Danish Pruthi
arXiv AI
Sep 16

Mapping U.S. Federal AI Governance Against Sector Vulnerability

The study evaluates 684 U.S. federal AI governance documents for how they address 14 sectors and 24 AI risks, measuring both breadth and depth of coverage. It finds that risks related to robustness, system security, and governance are more frequently and substantively discussed than socioeconomic, environmental, and emerging risks, and that sectors such as public administration, national security, information, and scientific services receive higher coverage than finance and healthcare. By comparing these coverage patterns with expert vulnerability assessments, the authors identify potential gaps in AI governance that could inform future policy and industry decisions.

By Ho Ting Hung, Angelica Chowdhury, James Teague, Simon Mylius, Spencer Michaels, Peter Slattery, Alexander Saeri, Neil Thompson
arXiv AI
Sep 23

Beyond Predictable Paths: Redefining AI Security Incident Reporting for Agents

The paper discusses the need to adapt incident reporting frameworks for AI agents, which are rapidly deployed and face unique security challenges. By comparing AI systems and agents and consulting 23 experts, the authors identify key reporting elements such as agent memory, autonomy levels, and tool usage. They also highlight open research questions, potential reporting weaknesses like data leakage, and outline privacy requirements for secure AI agent deployment.

By Anastasia Pustozerova, Eugene Bagdasarian, Luca Beurer-Kellner, Battista Biggio, Nico Ebert, David Filip, Marc Fischer, Heather Frase, David Hofer, Juliane Hoffmann, Daphne Ippolito, Somesh Jha, Sean McGregor, Esfandiar Mohammadi, Luca Nannini, Cristina Nita-Rotaru, Alina Oprea, Kevin Paeth, Andrew Paverd, Jonathan Petit, Andreas Rauber, Christian Riess, John Sotiropoulos, Andreas Wespi, Kathrin Grosse
arXiv AI
Sep 18

Reproducibility is not construct validity: LLM measurement of institutionally situated communication

The study shows that while large language model (LLM) annotations of stakeholder consultation submissions are highly reproducible (intraclass correlations > 0.99), they do not reliably capture the intended construct measured by structured survey responses. Divergence between LLM-inferred and survey measures varies by stakeholder group, with business associations expressing more AI risk concern in text than in surveys, and spatial autocorrelation indicates neighboring European countries share similar text-based stances. Despite these divergences, survey-reported concerns remain strongly linked to support for explainability across all levels of divergence.

By Veronika Batzdorfer (KIT), Carlo Romano Marcello Alessandro Santagiustina (ALMAnaCH, m\'edialab, Sciences Po)