arXiv AI By Bo Deng (Beihang University, Qwen DianJin Team, Alibaba Cloud Computing), Kang Zhou (Qwen DianJin Team, Alibaba Cloud Computing), Lifan Guo (Qwen DianJin Team, Alibaba Cloud Computing), Chongyang Tao (Beihang University), Xuanren Chen (Beihang University), Chenggang Xie (Beihang University), Renzhao Liang (Beihang University), Feng Chen (Qwen DianJin Team, Alibaba Cloud Computing), Chi Zhang (Qwen DianJin Team, Alibaba Cloud Computing)

FinEvo-Bench: A Longitudinal Benchmark for Self-Evolving Agents in Professional Financial Workflows

Read the original on arXiv AI →

arXiv:2608. 06144v1 Announce Type: new Abstract: Most agent benchmarks evaluate tasks independently and cannot measure whether experience from one task helps with later tasks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 18

FINSKILLOPS: A Self-Evolving Multi-Agent System for SEC Filing QA

FINSKILLOPS is a multi‑agent system designed to improve SEC filing question‑answering after deployment by creating reusable skills from evidence‑grounded failure diagnoses. It manages these skills through targeted validation, regression checks, negative controls, and versioned replacement or retirement, ensuring that new patches do not introduce regressions. Across six benchmarks, a frozen skill registry outperforms other systems, and in a 12‑round operational study the system reduced the non‑correct rate from 20.0% to 12.5% while promoting only six of 33 proposed skills.

By Yanzhang Ma, Zhenghan Tai, Hanwei Wu, Sizhe Guan, Jianliang Lei, Hailin He, Chaolong Jiang, Jijun Chi, Tung Sum Thomas Kwok, Bohuai Xiao, Jingrui Tian, Xinlu Wu, Xingao Zhan, Peng Lu, Muzhi Li, Yihong Wu, Liheng Ma, Sicheng Lyu, Tianshuo Yan, Junhao Zhu, Yaqian Xu, Lei Ding, Yufei Cui, Ziquan Liu, Boyu Han, Hengli Liu, Ling Zhou, Xinyu Wang