AssetOpsBench: Benchmarking AI Agents for Task Automation in Industrial Asset Operations and Maintenance
Related stories
DynAMO:Dynamic Asset Management Orchestration via Topological Multi-Agent Scheduling
arXiv:2606. 19382v1 Announce Type: cross Abstract: While LLM-powered agents offer end-to-end automation for industrial asset lifecycles, real-world Industry 4.
StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows
StartupBench is a new benchmark that evaluates general-purpose agents on end-to-end workflows derived from AI startup products that have proven market adoption. It translates real-world product workflows into deliverable-oriented tasks and assesses them with detailed rubrics. The study finds that even the best models complete only about 30% of these tasks, highlighting challenges such as complex instruction following and domain expertise.
AgenticDataBench: A Comprehensive Benchmark for Data Agents
arXiv:2607. 01647v1 Announce Type: cross Abstract: Data science aims to derive actionable insights from heterogeneous raw data, unlocking the value of the massive amounts of data generated in modern society.
StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows
StartupBench is a benchmark that evaluates general‑purpose agents on end‑to‑end workflows derived from real AI startup products that have proven market adoption. It translates these product workflows into deliverable‑oriented tasks and assesses them with detailed rubrics that capture complex requirements. Even the best current models complete only about 30% of the tasks, highlighting failures in instruction following and domain expertise.
Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields
arXiv:2606. 11042v1 Announce Type: new Abstract: Recent years have witnessed the rapid evolution of AI agents toward handling increasingly complex, real-world tasks.
AgentRx: Diagnosing AI Agent Failures from Execution Trajectories
arXiv:2602.02475v2 Announce Type: replace Abstract: AI agents often fail in ways that are difficult to localize because executions are probabilistic, long-horizon, multi-agent, and mediated by noisy...
Gypscie: A Cross-Platform AI Artifact Management System
arXiv:2604. 10311v2 Announce Type: replace Abstract: Artificial Intelligence (AI) models, encompassing both traditional machine learning (ML) and more advanced approaches such as deep learning and large language models (LLMs), play a central role in modern applications.
Characterizing Large Language Model Agentic Workflows: A Study on N8n Ecosystem
arXiv:2606. 29116v1 Announce Type: new Abstract: Large Language Models (LLMs) are rapidly being adopted in low-code and no-code automation platforms, where non-expert users design workflows that combine natural language understanding with external services and APIs.
Synthetic Scenario Generation for Evaluation of Industry 4.0 Agents
arXiv:2607. 22563v1 Announce Type: new Abstract: Industrial agent benchmarks require realistic evaluation scenarios that integrate telemetry, failure modes, maintenance records, and domain standards.
OS-Marathon: Benchmarking Computer-Use Agents on Vast-Horizon, Repetitive Tasks
OS-Marathon is a new benchmark that tests computer‑use agents on vast‑horizon, repetitive tasks, covering 100 tasks across five scenarios and ten domains. The study shows that current state‑of‑the‑art agents perform poorly on these tasks, and that simply decomposing workflows into subtasks does not solve the problem. Introducing a cost‑friendly personalization method called GraphDemo, which adapts agents from a single human demonstration, improves performance, highlighting the value of human guidance for these challenging tasks.
Ontology-supported AI Model and Dataset Management
The paper introduces an ontology-supported platform designed to facilitate the exchange, usage, and analysis of AI models and datasets. It addresses the need for effective management of AI assets in industrial settings by providing a structured framework that reduces semantic gaps. A real‑time critical systems use case demonstrates the platform’s practical utility.