arXiv AI

Benchmark-Based Comparative Assessment of Publicly Benchmarked Indian Foundation Models: A Capability and Evaluation-Maturity Framework

arXiv:2608. 11891v1 Announce Type: cross Abstract: Governments increasingly fund indigenous foundation models to strengthen national AI capability, digital sovereignty, and multilingual computing.

arXiv AI
Aug 14

RAIL: An Automatic Classifier of the Artificial Intelligence Readiness Level

arXiv:2608. 13428v1 Announce Type: new Abstract: Assessing the maturity of artificial intelligence technologies is essential for investment decisions, project management, and policy monitoring, yet the available readiness frameworks are heterogeneous and difficult to apply automatically: the adaptation of Technology Readiness Levels to AI lacks AI-specific gating criteria, the Machine Learning Technology Readiness Levels presuppose access to internal process artifacts, and AI/data readiness dimension models employ scales that resist direct comparison.

By Juan Irving Vasquez, Juan Terven, Laura-Ivoone Garay-Jimenez
arXiv AI
Aug 28

Thomson: Continual Learning of Frontier Models for SovereignAI

The paper introduces Thomson, a frontier AI model developed through continual learning on open-weight models, aiming to democratize access to high-performance AI. It argues that institutions with limited resources can achieve frontier-level performance by applying a modern mid- & post-training stack, preserving model plasticity and stability while minimizing high-impact interventions. Thomson demonstrates competitive performance across agentic tasks, safety, legal, tax, multilingualism, and large-scale deep research, exhibiting a distinctive π-shaped improvement pattern and effectively mitigating the forgetting problem seen in narrow domain adaptation.

By Shengzhuang Chen, Jerrod Parker, Yejin Bang, Andrew M. Bean, Nabeel Seedat, Stefan Winzeck, Daniil Glazko, Jannik Zgraggen, Fangyi Yu, Scott Arnott, Dietrich Trautmann, Luca Ciuffreda, Guglielmo Bonifazi, Davide Romano, Bradley Bell, Kirsty Fielding, Daniele Giofr\`e, Tom Zielund, Ipshita Chatterjee, Sneha Murthy Ghantasala, Manpreet Nanreh, John Scoville, Maciej Sakowicz, Wassim Seifeddine, Lukas Thede, Jonathan Richard Schwarz
arXiv AI
Aug 20

Global Index on Responsible AI 2026 : Conceptual Framework and Methodology

The Global Index on Responsible AI 2026 (GIRAI) 2nd Edition refines its predecessor by distinguishing between framework existence and implementation, expanding from three to five thematic areas, and adding granular variables for framework quality. It evaluates responsible AI governance across five dimensions—Inclusion and Diversity, Ethics and Sustainability, Labour and Skills, Trust and Safety, and Use of AI in Public Service—using 38 indicators organized into three pillars: AI Policy, CSO Engagement, and Enabling Conditions, plus a separate Use of Unacceptable Risk AI penalty. Data from 135 country-level researchers and secondary sources are normalized to a 100-point scale, weighted by pillar importance, and used to facilitate systematic cross‑national comparisons for policymakers, civil society, and AI developers.

By Fola Adeleke, Rachel Adams, Ayantola Alayande, Daniela Benavente, Ana Florido, Nicol\'as Grossman, Leah Junck
arXiv AI
Sep 18

A Unified Evaluation Framework for Trustworthy Large Language Models, Agentic AI, and Multimodal Systems

The paper introduces a unified evaluation framework for assessing the trustworthiness of large language models, agentic AI, and multimodal systems. It connects output-level, trajectory-level, and cross-modal assessments across eight dimensions—capability, robustness, safety, fairness, transparency, governance, oversight, and efficiency—while preserving system-specific metrics and providing uncertainty estimates. A meta-evaluation layer checks the validity, reliability, and reproducibility of the evaluation itself, and the framework aligns with governance standards and regulatory requirements.

By Shaina Raza, Ahmed Y. Radwan, Imran Liaquat, Kathryn Hume
arXiv Machine Learning
5d ago

One Capability or Many? Structural and Predictive Tests of Benchmark Validity Disagree About Economic Benchmarks for Frontier AI

The paper examines whether economic benchmarks used in frontier AI leaderboards measure a distinct capability or merely reflect general test-taking ability. Using a structural factor analysis and a predictive leave-one-benchmark-out test on a snapshot of 421 model configurations, the authors find that economic benchmarks do not form a separate factor but are better predicted by a multi‑factor representation than by a single general index, especially for linear learners. They argue that construct validity should be evaluated with both structural and predictive tests and provide a two‑test protocol along with data and code.

By Louis Yiven Zhu
arXiv AI
Aug 18

An Evaluation Framework for National AI Regulation

arXiv:2608. 15417v1 Announce Type: cross Abstract: Governments use laws, institutions, funding programs and nonbinding guidance to shape how AI is developed and used.

By Kaushik Sanjay Prabhakar, Tarun Adarsh R S, Amal Dhivyan Gregory, Sreeparvathy Sajeev, Utkarsh Tomar, Avyay M Casheekar
arXiv AI
Jul 3

PACE: A Proxy for Agentic Capability Evaluation

arXiv:2607. 02032v1 Announce Type: new Abstract: Evaluating LLM agents on benchmarks like SWE-Bench and GAIA can be expensive, time-consuming, and requires complex infrastructure.

By Yueqi Song, Lintang Sutawika, Jiarui Liu, Lindia Tjuatja, Jiayi Geng, Yunze Xiao, Daniel Lee, Aditya Bharat Soni, Vincent Lo, Xiang Yue, Graham Neubig