arXiv AI By Aryan Brar, Justin Du, Avery Lor, Kylie Seto, Eric Taylor

Retrieval-augmented generation vs. deterministic tax computation in multi-agent financial advisory: A 2x2 factorial experiment

Read the original on arXiv AI →

The study compares a custom capital gains calculation engine with a retrieval‑augmented generation (RAG) vector store of market advisory reports in a multi‑agent financial advisory system. A 2x2 factorial experiment showed that the tax‑optimization engine significantly reduced tax savings, while the RAG component had no significant effect. The RAG‑only condition yielded the highest tax savings, suggesting that pretrained language model knowledge may suffice for tax‑loss harvesting without specialized tooling.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 20

FinSkillBench: Evaluating AI Agents and Domain Skills for Investment Management

FinSkillBench is an evaluation suite that tests whether language model agents can use financial domain skills to solve investment management tasks across portfolio construction, risk management, and fundamental analysis. The benchmark contains 12 subtasks with 2,603 episodes, each providing point‑in‑time inputs, hidden ground truth, and a verifier. Experiments show that curated skill packages improve performance significantly, while self‑generated skills offer little benefit, indicating that reliable procedural skills are crucial for effective AI agents in this domain.

By Jermyn Zhen Yong Bek, Zhuang Qiang Bok, Zhongtian Sun
arXiv AI
Sep 17

EvolveTrade: Experience-Driven Policy Refinement for Self-Evolving LLM Trading Agents

EvolveTrade is a self‑evolving framework that treats the system prompt of a tool‑using LLM trading agent as a text‑parameterized policy. After each update interval, a Policy Agent revises this policy using accumulated decision traces and portfolio feedback while keeping the backbone LLM fixed, allowing the agent to refine its information‑acquisition and portfolio‑construction procedures over time. Experiments across multiple market regimes and two LLM backbones show that EvolveTrade often improves Sharpe Ratio and Cumulative Return over fixed‑policy baselines, with behavioral analyses indicating increased code‑mediated analysis and regime‑relevant computations.

By Sehee Kim, Yumin Choi, Minki Kang, Sung Ju Hwang