arXiv AI By Yanzhang Ma, Zhenghan Tai, Hanwei Wu, Sizhe Guan, Jianliang Lei, Hailin He, Chaolong Jiang, Jijun Chi, Tung Sum Thomas Kwok, Bohuai Xiao, Jingrui Tian, Xinlu Wu, Xingao Zhan, Peng Lu, Muzhi Li, Yihong Wu, Liheng Ma, Sicheng Lyu, Tianshuo Yan, Junhao Zhu, Yaqian Xu, Lei Ding, Yufei Cui, Ziquan Liu, Boyu Han, Hengli Liu, Ling Zhou, Xinyu Wang

FINSKILLOPS: A Self-Evolving Multi-Agent System for SEC Filing QA

Read the original on arXiv AI →

FINSKILLOPS is a multi‑agent system designed to improve SEC filing question‑answering after deployment by creating reusable skills from evidence‑grounded failure diagnoses. It manages these skills through targeted validation, regression checks, negative controls, and versioned replacement or retirement, ensuring that new patches do not introduce regressions. Across six benchmarks, a frozen skill registry outperforms other systems, and in a 12‑round operational study the system reduced the non‑correct rate from 20.0% to 12.5% while promoting only six of 33 proposed skills.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 7

FinEvo-Bench: A Longitudinal Benchmark for Self-Evolving Agents in Professional Financial Workflows

arXiv:2608. 06144v1 Announce Type: new Abstract: Most agent benchmarks evaluate tasks independently and cannot measure whether experience from one task helps with later tasks.

By Bo Deng (Beihang University, Qwen DianJin Team, Alibaba Cloud Computing), Kang Zhou (Qwen DianJin Team, Alibaba Cloud Computing), Lifan Guo (Qwen DianJin Team, Alibaba Cloud Computing), Chongyang Tao (Beihang University), Xuanren Chen (Beihang University), Chenggang Xie (Beihang University), Renzhao Liang (Beihang University), Feng Chen (Qwen DianJin Team, Alibaba Cloud Computing), Chi Zhang (Qwen DianJin Team, Alibaba Cloud Computing)
Hugging Face Trending Papers
Aug 18

TRUSS: Towards Task-Reliable and User-Safe Automated Agent Skill Generation

TRUSS is a framework that generates and verifies automated agent skills, ensuring they are both functionally effective and safe. It first checks functional claims against evidence and evaluates artifacts against nine safety properties, then tests admitted skills in a controlled environment to capture execution traces and identify failures. The approach achieves perfect precision and recall in vulnerability detection, significantly reduces attack success rates, and boosts task effectiveness and security rates in skill generation benchmarks.

arXiv Computation and Language
Aug 28

Towards Expert Financial QA via Self-Improving RAG

The paper introduces Self-Improving Retrieval-Augmented Generation (RAG), a framework that splits document question answering into Retrieval, Reasoning, and Judge agents coordinated by an orchestrator. When the Judge scores an answer below a dynamic threshold, the system retries with broader retrieval, more careful prompting, and relaxed acceptance criteria, achieving 86% oracle-guided accuracy on FinanceBench with a 36.4% Lazarus Rate. The approach logs every decision with confidence scores, providing audit trails needed for regulated financial applications.

By Junjie Xiong, Shawheen Ghezavat, Aum Hirpara
arXiv AI
Aug 19

TRUSS: Towards Task-Reliable and User-Safe Automated Agent Skill Generation

TRUSS is a framework that generates and verifies automated agent skills, ensuring they are both functionally effective and safe. It evaluates candidate skills against source evidence and nine safety properties, then tests them in a controlled environment to capture execution traces and identify failures. The system iteratively refines skills based on these results, achieving high precision in vulnerability detection and significantly improving task performance and security rates.

By Zhibo Zhang, Zhen Ouyang, Ling Shi, Kailong Wang