arXiv AI

RAIL: An Automatic Classifier of the Artificial Intelligence Readiness Level

arXiv:2608. 13428v1 Announce Type: new Abstract: Assessing the maturity of artificial intelligence technologies is essential for investment decisions, project management, and policy monitoring, yet the available readiness frameworks are heterogeneous and difficult to apply automatically: the adaptation of Technology Readiness Levels to AI lacks AI-specific gating criteria, the Machine Learning Technology Readiness Levels presuppose access to internal process artifacts, and AI/data readiness dimension models employ scales that resist direct comparison.

arXiv AI
Jul 7

Agentic and Generative AI for Open-Source Intelligence and Cyber Investigations: Taxonomy, Evaluation, Challenges, and Future Directions

arXiv:2607. 03233v1 Announce Type: cross Abstract: The rapid growth of publicly available digital information has rendered manual open-source intelligence (OSINT) analysis insufficient for modern intelligence, cybersecurity, and cyber investigation.

By Eduardo Almeida Palmieri, Mohamed Chahine Ghanem, Dipo Dunsin, Zubair Baig, Ed de Quincey, Kim-Kwang Raymond Choo
arXiv AI
Sep 18

A Unified Evaluation Framework for Trustworthy Large Language Models, Agentic AI, and Multimodal Systems

The paper introduces a unified evaluation framework for assessing the trustworthiness of large language models, agentic AI, and multimodal systems. It connects output-level, trajectory-level, and cross-modal assessments across eight dimensions—capability, robustness, safety, fairness, transparency, governance, oversight, and efficiency—while preserving system-specific metrics and providing uncertainty estimates. A meta-evaluation layer checks the validity, reliability, and reproducibility of the evaluation itself, and the framework aligns with governance standards and regulatory requirements.

By Shaina Raza, Ahmed Y. Radwan, Imran Liaquat, Kathryn Hume
arXiv AI
Sep 25

The Gold in Bias: Maturing the AI Design Process through Verification

The paper proposes rethinking bias in AI as a diagnostic tool rather than merely a flaw to be minimized. It introduces a multidimensional framework that examines bias across origin, lifecycle emergence, technical causes, and validation methods, covering 30 bias types, 16 verification methods, and 20 countermeasures for both traditional and generative AI. The authors present a hierarchical evidence framework distinguishing internal and external validity, and advocate for Ethics by Design principles to embed bias verification throughout the AI development lifecycle.

By Samira Maghool, Paolo Ceravolo
arXiv AI
Sep 15

Governing at Machine Speed: An Adaptive Intelligence Architecture for Real-Time AI Policy Enforcement

The paper "Governing at Machine Speed: An Adaptive Intelligence Architecture for Real-Time AI Policy Enforcement" highlights a gap in enterprise AI governance, where 78% of organizations lack auditable evidence of policy enforcement. It introduces AGIL, a five-layer adaptive governance architecture that uses machine learning for real-time detection, risk classification, sub-100ms policy enforcement, continuous attestation, and policy evolution. The authors argue that the failure is organizational and architectural, not technical, and call for future empirical validation of AGIL.

By Sandeep Bokkasam, B. Durgalakshmi
arXiv AI
Aug 24

Six misconceptions about large language models: A minimal model and diagnostic taxonomy

The article presents a minimal working model for large language model (LLM) systems, emphasizing four key distinctions—pretraining vs. deployment, distribution vs. samples, types of memory, and task competence vs. agency. Using this framework, it diagnoses six common misconceptions about LLMs (next‑token prediction, regression to the mean, training‑data regurgitation, model memory, alignment, and understanding), explaining what each misconception captures correctly, where it conflates distinctions, and the implications for evaluation, design, and governance. The model is applied to AI policy language, illustrating how policy can misrepresent these distinctions and offering a diagnostic toolkit to correct such errors.

By Zhicheng Lin