arXiv AI

Benchmarking Open-Weight Foundation Models for Global AI Technical Governance

arXiv:2606. 26099v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed in artificial intelligence (AI) governance analysis across national and international organisations.

arXiv AI
Aug 20

Global Index on Responsible AI 2026 : Conceptual Framework and Methodology

The Global Index on Responsible AI 2026 (GIRAI) 2nd Edition refines its predecessor by distinguishing between framework existence and implementation, expanding from three to five thematic areas, and adding granular variables for framework quality. It evaluates responsible AI governance across five dimensions—Inclusion and Diversity, Ethics and Sustainability, Labour and Skills, Trust and Safety, and Use of AI in Public Service—using 38 indicators organized into three pillars: AI Policy, CSO Engagement, and Enabling Conditions, plus a separate Use of Unacceptable Risk AI penalty. Data from 135 country-level researchers and secondary sources are normalized to a 100-point scale, weighted by pillar importance, and used to facilitate systematic cross‑national comparisons for policymakers, civil society, and AI developers.

By Fola Adeleke, Rachel Adams, Ayantola Alayande, Daniela Benavente, Ana Florido, Nicol\'as Grossman, Leah Junck
arXiv AI
Jul 17

Global Index on Responsible AI: 2026 Report

arXiv:2607. 14782v1 Announce Type: new Abstract: Grounded in human rights-based frameworks such as the UNESCO Recommendation on the Ethics of AI, the Global Index on Responsible AI (GIRAI) examines how countries translate responsible AI commitments into enforceable protections, institutional capacity, and redress mechanisms.

By Rachel Adams, Fola Adeleke, Ayantola Alayande, Selamawit Engida Abdella, Ana Florido, Nicol\'as Grossman, Leah Junck
arXiv AI
Aug 20

Computational Orientalism: Measuring Structural Discourse Bias in Large Language Models Using the Middle East Cultural Sensitivity Score (MECSS)

The paper introduces the Middle East Cultural Sensitivity Score (MECSS) to quantify Orientalist bias in large language models, converting Said’s seven Orientalist operations into measurable dimensions. Using 280 conversations, it finds that GPT‑4 and Falcon3‑7B‑Instruct systematically reproduce Orientalist patterns, with Falcon scoring higher despite being regionally built. The study highlights that geographic origin alone does not mitigate bias and identifies a new failure mode, "Said‑washing," present in 87.9% of GPT‑4 interactions.

By Maha Shahid
arXiv AI
Aug 24

Six misconceptions about large language models: A minimal model and diagnostic taxonomy

The article presents a minimal working model for large language model (LLM) systems, emphasizing four key distinctions—pretraining vs. deployment, distribution vs. samples, types of memory, and task competence vs. agency. Using this framework, it diagnoses six common misconceptions about LLMs (next‑token prediction, regression to the mean, training‑data regurgitation, model memory, alignment, and understanding), explaining what each misconception captures correctly, where it conflates distinctions, and the implications for evaluation, design, and governance. The model is applied to AI policy language, illustrating how policy can misrepresent these distinctions and offering a diagnostic toolkit to correct such errors.

By Zhicheng Lin
arXiv AI
Jul 1

RCTs for Frontier AI Governance: Methodological Challenges and Solutions for Human Uplift Studies

arXiv:2603. 11001v3 Announce Type: replace-cross Abstract: Human uplift studies, or studies that measure the effects of AI access on human performance via randomized controlled trials (RCT) or similar methodologies, increasingly inform frontier AI governance and deployment decisions.

By Patricia Paskov, Kevin Wei, Shen Zhou Hong, Dan Bateyko, Xavier Roberts-Gaal, Carson Ezell, Gailius Praninskas, Valerie Chen, Umang Bhatt, Ella Guest
arXiv AI
Sep 16

Toward Governance-Aware Autonomous GIS: A Narrative Review of Ethical and Privacy Risks in LLM-Enabled GeoAI

The paper reviews ethical and privacy risks of large language model (LLM)–enabled geospatial artificial intelligence (GeoAI), identifying eight recurring issues such as data provenance, spatial privacy, algorithmic bias, and technical risks. It evaluates current responses, noting many remain largely unaddressed or conceptual, and proposes a governance‑aware architecture with enforceable controls illustrated by a flood‑response routing scenario. The authors call for empirical validation, spatially specific interpretability tools, and workforce training to address these emerging risks.

By Maya Subramanian, Devika Jain
arXiv AI
Aug 28

Thomson: Continual Learning of Frontier Models for SovereignAI

The paper introduces Thomson, a frontier AI model developed through continual learning on open-weight models, aiming to democratize access to high-performance AI. It argues that institutions with limited resources can achieve frontier-level performance by applying a modern mid- & post-training stack, preserving model plasticity and stability while minimizing high-impact interventions. Thomson demonstrates competitive performance across agentic tasks, safety, legal, tax, multilingualism, and large-scale deep research, exhibiting a distinctive π-shaped improvement pattern and effectively mitigating the forgetting problem seen in narrow domain adaptation.

By Shengzhuang Chen, Jerrod Parker, Yejin Bang, Andrew M. Bean, Nabeel Seedat, Stefan Winzeck, Daniil Glazko, Jannik Zgraggen, Fangyi Yu, Scott Arnott, Dietrich Trautmann, Luca Ciuffreda, Guglielmo Bonifazi, Davide Romano, Bradley Bell, Kirsty Fielding, Daniele Giofr\`e, Tom Zielund, Ipshita Chatterjee, Sneha Murthy Ghantasala, Manpreet Nanreh, John Scoville, Maciej Sakowicz, Wassim Seifeddine, Lukas Thede, Jonathan Richard Schwarz
arXiv AI
Aug 7

Who Gets Access? Global Region and Academic Status Bias in AI-Generated Academic Gatekeeping Scenarios

arXiv:2608. 05178v1 Announce Type: cross Abstract: Equitable access to scientific knowledge often depends on informal gatekeeping decisions, particularly when resources such as paywalled articles, datasets, or professional materials such as curriculum vitae (CV) must be shared selectively.

By Nouar AlDahoul, Hezerul Abdul Karim, Myles Joshua Toledo Tan
arXiv AI
Aug 14

RAIL: An Automatic Classifier of the Artificial Intelligence Readiness Level

arXiv:2608. 13428v1 Announce Type: new Abstract: Assessing the maturity of artificial intelligence technologies is essential for investment decisions, project management, and policy monitoring, yet the available readiness frameworks are heterogeneous and difficult to apply automatically: the adaptation of Technology Readiness Levels to AI lacks AI-specific gating criteria, the Machine Learning Technology Readiness Levels presuppose access to internal process artifacts, and AI/data readiness dimension models employ scales that resist direct comparison.

By Juan Irving Vasquez, Juan Terven, Laura-Ivoone Garay-Jimenez