arXiv AI

Design Methodology and Performance Trade-offs Management for Distributed and Compound AI Systems

arXiv:2606. 14350v1 Announce Type: cross Abstract: Artificial Intelligence (AI) systems must typically satisfy service-level objectives including accuracy, latency, and cost.

arXiv AI
Sep 7

Atlas: Optimizing Deployment of Compound AI Workflows on Heterogeneous Clusters

Atlas is a framework that optimizes the deployment of compound AI workflows on heterogeneous clusters by selecting execution plans that satisfy service level objectives (SLOs). It introduces MAP, a Markovian Accuracy Predictor, which estimates configuration accuracy using local conditional accuracy transitions between adjacent workflow stages, avoiding exhaustive end‑to‑end profiling. Atlas formulates plan selection as a mixed‑integer linear program, achieving near‑oracle accuracy while reducing deployment cost by up to 42% and profiling cost by up to 2.6×.

By Milos Gravara, Andrija Stanisic, Stefan Nastic
arXiv AI
Jun 15

PLAIground: SLO-Driven Runtime Model Selection for Compound AI Systems in the Edge-Cloud-Space Continuum

arXiv:2606. 14356v1 Announce Type: cross Abstract: Applications in the 3D Computing Continuum, which unifies edge, cloud, and space, require combining multiple AI tasks such as object detection, time-series analytics, and natural language processing into Compound AI systems.

By Milos Gravara, Cynthia Marcelino, Andrija Stanisic, Stefan Nastic
arXiv AI
Aug 28

Benchmarking AI Agents for Hardware Design Automation via MCP Tool Calling

The paper investigates whether locally deployed large language models can automate hardware design workflows that involve repetitive, dependency-ordered operations using specialized tools. A Model Context Protocol (MCP) server is created to emulate a proprietary hardware design tool, and a benchmark tests single and multi-step edits, invalid requests, misspelled prompts, and multi-server contexts. Seven open-source models are evaluated across different pipeline choices, revealing that strong models can nearly fully cover expected calls, but reliability hinges on task structure and agent configuration, with comprehensive tool descriptions reducing failures and multi-agent setups aiding weaker models at the cost of extra calls.

By Leonardo Liparulo, Francesco Pierri
arXiv AI
Jul 7

Gypscie: A Cross-Platform AI Artifact Management System

arXiv:2604. 10311v2 Announce Type: replace Abstract: Artificial Intelligence (AI) models, encompassing both traditional machine learning (ML) and more advanced approaches such as deep learning and large language models (LLMs), play a central role in modern applications.

By Fabio Porto, Eduardo Ogasawara, Gabriela Moraes Botaro, Julia Neumann Bastos, Augusto Fonseca, Esther Pacitti, Patrick Valduriez
arXiv AI
Sep 18

Architectural Design, Not Only Model Intelligence, Governs Multi-Agent LLM Performance

The paper argues that the architecture of multi‑agent large language model (LLM) frameworks, rather than just the intelligence of the underlying models, largely determines system performance. It introduces a taxonomy of architectural dimensions—such as orchestration, memory, planning interfaces, specialization, and communication topology—and presents MAFBench, a unified evaluation suite. An empirical study across nine frameworks, keeping the LLM constant, reveals six design principles and shows that choices like orchestration and communication topology can dramatically affect latency, accuracy, and coordination success.

By Abdelghny Orogat, Ana Rostam, Essam Mansour
arXiv AI
Jul 14

A Formal Hierarchical Architecture for Agentic Orchestration with Stack-Based Execution and Lazy Discovery

arXiv:2607. 11138v1 Announce Type: new Abstract: The rapid expansion of capabilities in Large Language Model (LLM) agents has exposed a critical architectural bottleneck: when agents are given access to a flat, monolithic registry of tools, the model must evaluate hundreds or thousands of options simultaneously.

By Prashant Devadiga, Abhishek, Adithya Mishra, Alok Singh, Amisha Sinha, Asit Desai, Gaurang Dahad, Harshit Bhushan, Mandati Pramod Reddy, Prakhar Gupta, Rupesh Patil, Siddhi Behere
arXiv AI
Sep 1

SimCRAFT: Distilling Remote Sensing Agents via Synthetic Trajectories and Contextual Retrieval-Augmented Fine-Tuning

SimCRAFT is a model‑agnostic framework that distills remote sensing orchestration into a compact 7B‑scale model. It creates a large, constraint‑validated workflow planning corpus (SimRS‑14k) using a multi‑agent synthesis engine and a Mock Execution Engine, then fine‑tunes the model with Contextual Retrieval‑Augmented Fine‑Tuning (CRAFT) to reason analogically. Experiments show SimCRAFT‑7B outperforms open‑weight LLMs and rivals advanced closed‑source models, providing a lightweight, efficient baseline for autonomous remote sensing deployment.

By Haoran Wang, Jing Yao, Xu Yang, Zeqing Wang, Yang Zhang, Pedram Ghamisi, Zhengchao Chen
arXiv Machine Learning
Jun 9

OpenCompass: A Universal Evaluation Platform for Large Language Models

arXiv:2605. 19276v3 Announce Type: replace-cross Abstract: In recent years, the field of artificial intelligence has undergone a paradigm shift from task-specific small-scale models to general-purpose large language models (LLMs).

By Maosong Cao, Kai Chen, Haodong Duan, Yixiao Fang, Zhiwei Fei, Tong Gao, Ge Jiaye, Mo Li, Hongwei Liu, Junnan Liu, Yuan Liu, Chengqi Lyu, Han Lyu, Ningsheng Ma, Zerun Ma, Yu Sun, Zhiyong Wu, Linchen Xiao, Zhuozhi Xiong, Jun Xu, Haochen Ye, Zhaohui Yu, Yike Yuan, Songyang Zhang, Yufeng Zhao, Fengzhe Zhou, Peiheng Zhou, Dongsheng Zhu, Lin Zhu, Jingming Zhuo