arXiv:2607. 17963v1 Announce Type: new Abstract: Ontology extension refers to the process of enriching an existing ontology in response to emerging requirements, making it more complete.
By Anna Sofia Lippolis, Mohammad Javad Saeedizade, Stefan Schmid, Simon Blattner, Robin Keskis\"arkk\"a, Aldo Gangemi, Eva Blomqvist, Andrea Giovanni Nuzzolese
arXiv:2608. 04921v1 Announce Type: cross Abstract: As AI systems become increasingly integrated into diverse interfaces and applications, model-centric audits are insufficient to address risks arising from interactions among system components and deployment environments.
By Leah Davis, Dominic Martin, AJung Moon
arXiv:2606. 24619v1 Announce Type: new Abstract: Competency Questions (CQs) are the central component of CQ-verification, an established process in which an ontology is evaluated against a set of natural language questions to determine whether the intended purpose of the ontology has been properly modelled.
By Anna Sofia Lippolis, Mohammad Javad Saeedizade, Robin Keskis\"arkk\"a, Aldo Gangemi, Eva Blomqvist, Andrea Giovanni Nuzzolese
The paper introduces Constraint-Driven Context Engineering (CDCE), a design approach that treats domain constraints as primary drivers for creating AI system interfaces. CDCE identifies, characterises, and operationalises constraints to determine necessary context assets and their representations, improving the quality and domain appropriateness of AI-generated solutions. A comparative multiple‑case study across education, healthcare, and finance demonstrates CDCE’s applicability and shows how constraint characteristics shape the resulting interfaces.
By Xiwei Xu, Chen Wang, Mengmeng Yang, Yipeng Zhang, Jacky Jiang, Suyu Ma, Youyang Qu, Ming Ding, Liming Zhu
The paper argues that large language models (LLMs) are evaluated too narrowly, focusing on isolated technical metrics rather than holistic, developmental, and societal aspects. It proposes a diagnostic ontology that links evaluation dimensions to the LLM training pipeline, turning evaluation into a root‑cause analysis tool. The authors introduce an anthropomorphic framework—IQ, PQ, EQ, and VQ—to assess LLM capabilities, operationalize it with a modular architecture, and validate it through meta‑analysis of over 200 benchmarks, outlining key challenges and future directions.
By Jun Wang, Ninglun Gu, Kailai Zhang, Pengyong Li, Yelun Bao, Jin Yang, Xu Yin, Liwei Liu, Zijiao Zhang, Yihuan Liu, Gary G. Yen, Junchi Yan
The paper introduces a pipeline that automatically creates ontology‑grounded multiple‑choice question benchmarks for evaluating large language models (LLMs) on logical reasoning tasks in scientific AI. By using OWL 2 ontologies, correct answers are guaranteed by design and distractors are generated and formally verified as incorrect through an OWL reasoner. Experiments on three ontologies—Pizza, PMDco, and DOID—yielded 112, 2,491, and 15,216 MCQs, respectively, with high natural‑language quality and challenging zero‑shot performance for six LLMs.
By Nishtha N. Vaidya, Stephan Grimm, Thomas Hubauer, Thomas A. Runkler
The paper introduces an ontology-supported platform designed to facilitate the exchange, usage, and analysis of AI models and datasets. It addresses the need for effective management of AI assets in industrial settings by providing a structured framework that reduces semantic gaps. A real‑time critical systems use case demonstrates the platform’s practical utility.
By Jan Novacek, Ali Ahari, Tobias M\"uller, Sebastian Reiter, Alexander Viehl, Oliver Bringmann
arXiv:2606. 04037v1 Announce Type: new Abstract: Pre-deployment verification of enterprise artificial intelligence (AI) agents remains a critical gap between large language model (LLM) capability benchmarking and production deployment.
By Thanh Luong Tuan, Abhijit Sanyal
arXiv:2609.21841v1 Announce Type: new
Abstract: Frontier language models now produce professional deliverables that expert graders judge to match human work on a substantial share of economically val...
By Abbas Raza Ali, Muhammad Ajmal Siddiqui, Moona Zahid
arXiv:2607. 14673v1 Announce Type: new Abstract: Evaluations (Evals) are a deployment bottleneck for real-world AI applications: public benchmarks rarely match a team's users, context, or policies, and human review is often tedious to scale.
By Leanne Tan, Rohan Jaggi, Shaun Khoo, Roy Ka-Wei Lee
arXiv:2607. 05638v1 Announce Type: cross Abstract: Teams deploying large language models in business contexts need evaluation systems, yet most treat evaluation as static model selection: run benchmarks, rank models, deploy the winner.
By Kenneth Benavides, Josh Fleischer, Danti Chen
arXiv:2605. 22093v3 Announce Type: replace Abstract: Knowledge graphs have become the primary vehicle for data integration and are critical to the success of modern AI, but the diversity of KG modelling practices, from lightweight vocabularies to richly axiomatised ontologies, makes integration and reuse expensive and brittle.
By Enrico Daga, Valentina Tamma, Terry Payne