arXiv AI By Milos Gravara, Andrija Stanisic, Stefan Nastic

Design Methodology and Performance Trade-offs Management for Distributed and Compound AI Systems

Read the original on arXiv AI →

arXiv:2606. 14350v1 Announce Type: cross Abstract: Artificial Intelligence (AI) systems must typically satisfy service-level objectives including accuracy, latency, and cost.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 7

Atlas: Optimizing Deployment of Compound AI Workflows on Heterogeneous Clusters

Atlas is a framework that optimizes the deployment of compound AI workflows on heterogeneous clusters by selecting execution plans that satisfy service level objectives (SLOs). It introduces MAP, a Markovian Accuracy Predictor, which estimates configuration accuracy using local conditional accuracy transitions between adjacent workflow stages, avoiding exhaustive end‑to‑end profiling. Atlas formulates plan selection as a mixed‑integer linear program, achieving near‑oracle accuracy while reducing deployment cost by up to 42% and profiling cost by up to 2.6×.

By Milos Gravara, Andrija Stanisic, Stefan Nastic
arXiv AI
Jun 15

PLAIground: SLO-Driven Runtime Model Selection for Compound AI Systems in the Edge-Cloud-Space Continuum

arXiv:2606. 14356v1 Announce Type: cross Abstract: Applications in the 3D Computing Continuum, which unifies edge, cloud, and space, require combining multiple AI tasks such as object detection, time-series analytics, and natural language processing into Compound AI systems.

By Milos Gravara, Cynthia Marcelino, Andrija Stanisic, Stefan Nastic
arXiv AI
Aug 28

Benchmarking AI Agents for Hardware Design Automation via MCP Tool Calling

The paper investigates whether locally deployed large language models can automate hardware design workflows that involve repetitive, dependency-ordered operations using specialized tools. A Model Context Protocol (MCP) server is created to emulate a proprietary hardware design tool, and a benchmark tests single and multi-step edits, invalid requests, misspelled prompts, and multi-server contexts. Seven open-source models are evaluated across different pipeline choices, revealing that strong models can nearly fully cover expected calls, but reliability hinges on task structure and agent configuration, with comprehensive tool descriptions reducing failures and multi-agent setups aiding weaker models at the cost of extra calls.

By Leonardo Liparulo, Francesco Pierri