arXiv AI By Qi Cheng, Shengyu Chen, Wei Cheng, Yiqun Xie, Haoyu Wang, Haifeng Chen, Xiaowei Jia

RA-MoWE: Workflow-Affinity Embeddings for Query Clustering and Agentic Workflow Generation

Read the original on arXiv AI →

RA-MoWE introduces workflow‑affinity embeddings to cluster queries and guide the creation of reusable expert workflows for large language models. Each embedding captures how well a set of reference workflows solves a query, revealing common reasoning strategies. The framework uses cluster embeddings to initialize and refine specialized workflows, and an encoder predicts embeddings from query text, enabling efficient expert selection without executing reference workflows.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
1d ago

OOPMAS: Object-Oriented Multi-Agent Systems for Query-Level Workflow Generation

OOPMAS introduces a training‑free framework that generates both the agent set and the coordination workflow at the granularity of individual queries. Agents are defined as object‑oriented class definitions with dedicated roles, tools, and persistent state, while workflows are expressed as executable main functions over these agent objects. A dynamic skill library accumulates structured lessons from execution feedback across optimization rounds, enabling in‑context improvement without any gradient updates or fine‑tuning, and achieves 89.6% accuracy on a mixed‑task benchmark, outperforming the strongest baseline by 18.1 percentage points.

By Qi Cheng, Shengyu Chen, Wei Cheng, Yiqun Xie, Xiaowei Jia, Haoyu Wang, Haifeng Chen
arXiv Computation and Language
Aug 25

LongWoF-Bench: Evaluating EvoMap Genes for Verifiable Long-Workflow Tasks

LongWoF-Bench is a new benchmark of 778 machine‑verifiable long‑workflow tasks spanning code generation, agent‑environment synthesis, mathematical reasoning, and rule following. The study shows that EvoMap Genes—structured representations of verifier‑confirmed execution trajectories—outperform the Skill baseline by 8.7–15.5 percentage points across seven models, and for Claude Opus they enable 39 additional task completions while cutting token consumption by 9.9%. The results demonstrate that verified execution experience can be externalized and reused, improving long‑workflow completion without repeatedly discovering new strategies.

By Xiao Zhang, Qumeng Sun, Jihao Li, Yiming Ren, Xiang Liu, Haoyang Zhang, Junjie Wang