arXiv:2508. 16181v2 Announce Type: replace-cross Abstract: Cross-organizational collaboration in Model-Based Systems Engineering (MBSE) faces many challenges in achieving semantic alignment across independently developed system models.
By Zirui Li, Stephan Husung, Haoze Wang
arXiv:2608. 04921v1 Announce Type: cross Abstract: As AI systems become increasingly integrated into diverse interfaces and applications, model-centric audits are insufficient to address risks arising from interactions among system components and deployment environments.
By Leah Davis, Dominic Martin, AJung Moon
arXiv:2607. 14659v1 Announce Type: cross Abstract: Interoperability between heterogeneous modeling tools remains a significant challenge in Model-Driven Engineering (MDE), particularly in the automotive domain where multiple modeling languages, as well as defacto standard proprietary and open-source tools coexist.
By Nenad Petrovic, Jiajie Zhang, Vahid Zolfaghari, Alois Knoll
Large language model (LLM) agents are increasingly evaluated on their ability to use tools, plan multi-step tasks, coordinate with other agents, and operate over extended horizons. Reported benchmark gains often obscure recurring failure modes documented across otherwise unrelated evaluation efforts.
arXiv:2609. 04377v1 Announce Type: new Abstract: Enterprise AI deployments fail not from model inadequacy, but because organizations lack a structured substrate encoding how they decide, negotiate, and execute.
By Fabricio C. Avini, Guilherme Trez
The paper introduces a unified evaluation framework for assessing the trustworthiness of large language models, agentic AI, and multimodal systems. It connects output-level, trajectory-level, and cross-modal assessments across eight dimensions—capability, robustness, safety, fairness, transparency, governance, oversight, and efficiency—while preserving system-specific metrics and providing uncertainty estimates. A meta-evaluation layer checks the validity, reliability, and reproducibility of the evaluation itself, and the framework aligns with governance standards and regulatory requirements.
By Shaina Raza, Ahmed Y. Radwan, Imran Liaquat, Kathryn Hume
arXiv:2606. 06535v1 Announce Type: cross Abstract: Context.
By Faezeh Amou Najafabad, Markus Haug, Keerthiga Rajenthiram, Justus Bogner, Ilias Gerostathopoulos
arXiv:2607. 05775v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly evaluated on their ability to use tools, plan multi-step tasks, coordinate with other agents, and operate over extended horizons.
By Wael Albayaydh, Rui Zhao, Ivan Flechais
arXiv:2602.24055v5 Announce Type: replace
Abstract: This study proposes CIRCLE, a six-stage, lifecycle-based framework to bridge the reality gap between model-centric performance metrics and AI syste...
By Reva Schwartz, Carina Westling, Morgan Briggs, Marzieh Fadaee, Isar Nejadgholi, Matthew Holmes, Fariza Rashid, Maya Carlyle, Afaf Ta\"ik, Kyra Wilson, Peter Douglas, Theodora Skeadas, Gabriella Waters, Rumman Chowdhury, Thiago Lacerda
arXiv:2606. 29116v1 Announce Type: new Abstract: Large Language Models (LLMs) are rapidly being adopted in low-code and no-code automation platforms, where non-expert users design workflows that combine natural language understanding with external services and APIs.
By Yutian Tang, Yuming Zhou, Huaming Chen
arXiv:2607. 17331v1 Announce Type: new Abstract: Enterprise Resource Planning (ERP) systems record transactions reliably but still delegate almost all operational decision-making to human specialists, because classical rule-based automation cannot reason about exceptions and monolithic AI assistants degrade when asked to coordinate across functional boundaries.
By Zhihao Liu, Tianyu Wang, Xi Vincent Wang, Lihui Wang
arXiv:2608. 03413v1 Announce Type: new Abstract: As artificial intelligence (AI) continues to evolve and mature, recent AI practices have moved beyond large language models (LLMs) and text or image generation tasks, increasingly integrating tools, agents, and harnesses to solve real business and industrial problems.
By Zuojun Max Shen, Yuan Qu, Pujun Zhang, Anbang Liu, Yunhao Liang