FinProBench: Evaluating Financial AI Agents with Role-Grounded Rubrics Derived from Professional Deliverables
arXiv:2608. 04077v1 Announce Type: new Abstract: Evaluating financial AI agents requires criteria aligned with real professional work.
arXiv:2608. 06144v1 Announce Type: new Abstract: Most agent benchmarks evaluate tasks independently and cannot measure whether experience from one task helps with later tasks.
arXiv:2608. 04077v1 Announce Type: new Abstract: Evaluating financial AI agents requires criteria aligned with real professional work.
arXiv:2607. 05297v1 Announce Type: new Abstract: Recent LLM agents tackle increasingly long-horizon, open-ended tasks, and external skills, reusable procedural knowledge supplied to the agent, further extend this capability.
FINSKILLOPS is a multi‑agent system designed to improve SEC filing question‑answering after deployment by creating reusable skills from evidence‑grounded failure diagnoses. It manages these skills through targeted validation, regression checks, negative controls, and versioned replacement or retirement, ensuring that new patches do not introduce regressions. Across six benchmarks, a frozen skill registry outperforms other systems, and in a 12‑round operational study the system reduced the non‑correct rate from 20.0% to 12.5% while promoting only six of 33 proposed skills.
arXiv:2608. 02636v1 Announce Type: cross Abstract: Self-evolving skill systems promise to improve agents by turning execution feedback into persistent skill updates without changing the underlying model.
arXiv:2606. 14239v1 Announce Type: new Abstract: Agent skills are structured procedural packages that guide frozen LLM agents in specialized workflows.
arXiv:2608. 03764v1 Announce Type: new Abstract: Agent self-evolution updates an agent's persistent state from prior experience and reuses it to solve related tasks more effectively.
arXiv:2607. 25891v1 Announce Type: new Abstract: Evaluating AI agents in interactive environments is hindered by fragmented tasks, scaffolds, verifiers, and scoring rules.
arXiv:2605. 17554v2 Announce Type: replace Abstract: Frontier deep research agents (DRAs) plan a research task, synthesize across documents, and return a structured deliverable on demand.
arXiv:2609.24663v1 Announce Type: new Abstract: Self-evolving agents convert interaction feedback into persistent artifacts, such as memories or skills, which in turn guide subsequent decisions. As t...
arXiv:2609.36746v1 Announce Type: new Abstract: Agent skills provide a lightweight mechanism for self-evolving agents to accumulate reusable procedural knowledge without updating model parameters. Ho...
FinSkillBench is an evaluation suite that tests whether language model agents can use financial domain skills to solve investment management tasks across portfolio construction, risk management, and fundamental analysis. The benchmark contains 12 subtasks with 2,603 episodes, each providing point‑in‑time inputs, hidden ground truth, and a verifier. Experiments show that curated skill packages improve performance significantly, while self‑generated skills offer little benefit, indicating that reliable procedural skills are crucial for effective AI agents in this domain.
Self-evolving agents convert interaction feedback into persistent artifacts, such as memories or skills, which in turn guide subsequent decisions. As these artifacts are iteratively updated throughout...