arXiv AI By Xuehang Guo, Haoyu Wang, Shengyu Chen, Zach Chen, Wei Cheng, Qingyun Wang, Haifeng Chen

Pay for the Fault, Not the Flow: Label-Free In-Flow Multi-Agent Workflow Optimization

Read the original on arXiv AI →

The paper introduces InFlowOp, a label‑free optimization framework that assigns costs to each decision in a multi‑agent workflow, balancing agent competence against execution time. It determines task granularity and agent assignment before execution and corrects faults during execution using the same cost metric. The authors also present Braid, a benchmark for multi‑agent coordination, and show that InFlowOp outperforms single‑agent baselines by up to 11.97% across various domains.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
2d ago

It Takes Workflows to Evolve Better Workflows

arXiv:2610.01026v1 Announce Type: cross Abstract: Tackling complex real-world tasks can exceed the capabilities of a single large language model (LLM), motivating the use of multi-agent workflows tha...

By Xuehang Guo, Haoyu Wang, Haifeng Chen, Yangyi Chen, Zhenhailong Wang, Qingyun Wang
arXiv AI
Sep 17

Designing Agentic AI Workflow Portfolios under Imperfect Selection and Compute Cost

The paper investigates how to design portfolios of agentic AI workflows that vary in reasoning strategy, verification structure, and compute cost. It proposes a portfolio-and-selector framework where multiple workflow executions are run and the best output is chosen, balancing additional compute with potential gains in accuracy. The authors develop exact and approximate optimization methods, evaluate them on three datasets, and show modest improvements over the best single workflow.

By Mojtaba Abdolmaleki, Stefanus Jasin, Boyu Wang
arXiv AI
Jul 23

In-the-Flow Agentic System Optimization for Effective Planning and Tool Use

arXiv:2510. 05592v2 Announce Type: replace Abstract: Outcome-driven reinforcement learning has advanced reasoning in large language models (LLMs), but prevailing tool-augmented approaches train a single, monolithic policy that interleaves thoughts and tool calls under full context; this scales poorly with long horizons and diverse tools and generalizes weakly to new scenarios.

By Zhuofeng Li, Haoxiang Zhang, Seungju Han, Sheng Liu, Jianwen Xie, Yu Zhang, Yejin Choi, James Zou, Pan Lu