arXiv AI By Julie Krugler Hollek, Michael Zargham, Mala Kumar

Ontology-Based Contextual AI Evaluations (OB-CAIE) Methodology

Read the original on arXiv AI →

The Ontology-Based Contextual AI Evaluations (OB-CAIE) methodology introduces a structured approach to AI evaluation by defining clear testing coverage and balancing human expertise with automation. It employs two ontologies—the Domain‑Specific Ontology (DSO) outlining what is tested, and the Evaluation Process Ontology (EPO) detailing how it is tested—to create a tractable problem space that can be applied to single or multiple AI evaluations. OB‑CAIE enables traceable, visualizable failure points and incorporates human judgment at scientifically grounded junctures where machine input is insufficient.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 24

When CQs Go Wrong: Challenges in CQ Verification with OE-Assist

arXiv:2606. 24619v1 Announce Type: new Abstract: Competency Questions (CQs) are the central component of CQ-verification, an established process in which an ontology is evaluated against a set of natural language questions to determine whether the intended purpose of the ontology has been properly modelled.

By Anna Sofia Lippolis, Mohammad Javad Saeedizade, Robin Keskis\"arkk\"a, Aldo Gangemi, Eva Blomqvist, Andrea Giovanni Nuzzolese
arXiv AI
Sep 24

Constraint-Driven Context Engineering: Designing Domain Interfaces for AI Systems

The paper introduces Constraint-Driven Context Engineering (CDCE), a design approach that treats domain constraints as primary drivers for creating AI system interfaces. CDCE identifies, characterises, and operationalises constraints to determine necessary context assets and their representations, improving the quality and domain appropriateness of AI-generated solutions. A comparative multiple‑case study across education, healthcare, and finance demonstrates CDCE’s applicability and shows how constraint characteristics shape the resulting interfaces.

By Xiwei Xu, Chen Wang, Mengmeng Yang, Yipeng Zhang, Jacky Jiang, Suyu Ma, Youyang Qu, Ming Ding, Liming Zhu
arXiv AI
Aug 25

Beyond Benchmarks: LLM Evaluation with an Anthropomorphic and Lifecycle-oriented Roadmap

The paper argues that large language models (LLMs) are evaluated too narrowly, focusing on isolated technical metrics rather than holistic, developmental, and societal aspects. It proposes a diagnostic ontology that links evaluation dimensions to the LLM training pipeline, turning evaluation into a root‑cause analysis tool. The authors introduce an anthropomorphic framework—IQ, PQ, EQ, and VQ—to assess LLM capabilities, operationalize it with a modular architecture, and validate it through meta‑analysis of over 200 benchmarks, outlining key challenges and future directions.

By Jun Wang, Ninglun Gu, Kailai Zhang, Pengyong Li, Yelun Bao, Jin Yang, Xu Yin, Liwei Liu, Zijiao Zhang, Yihuan Liu, Gary G. Yen, Junchi Yan
arXiv AI
2d ago

Ontology-Grounded, Reasoner-Verified Benchmarks for Evaluating LLM Reasoning in Scientific AI

The paper introduces a pipeline that automatically creates ontology‑grounded multiple‑choice question benchmarks for evaluating large language models (LLMs) on logical reasoning tasks in scientific AI. By using OWL 2 ontologies, correct answers are guaranteed by design and distractors are generated and formally verified as incorrect through an OWL reasoner. Experiments on three ontologies—Pizza, PMDco, and DOID—yielded 112, 2,491, and 15,216 MCQs, respectively, with high natural‑language quality and challenging zero‑shot performance for six LLMs.

By Nishtha N. Vaidya, Stephan Grimm, Thomas Hubauer, Thomas A. Runkler