arXiv Computation and Language

Fund2Persona: A Framework for Building and Refining Financial Advisor Personas from Fund Disclosure Data

Hugging Face Trending Papers
Sep 2

The Analyst in the Prompt: Role, Retrieval, and Memory Biases in LLM Financial Analysis

The paper investigates how user context—such as memory, profiles, and role prompts—affects large language models’ financial analysis. By testing 3,575 SEC filings across twelve LLMs, the study distinguishes between evidence selection and interpretation, finding that interpretation under different roles drives most user-context spillover. Two mitigation strategies—using a user profile instead of an assistant role and separating evidence-based from personalized outputs—reduce but do not eliminate this spillover, with effectiveness varying by model.

arXiv Computation and Language
Sep 4

The Analyst in the Prompt: Role, Retrieval, and Memory Biases in LLM Financial Analysis

The study examines how user context—such as memory, profiles, and role prompts—affects Large Language Models’ (LLMs) financial analysis. Using 3,575 SEC filings and twelve LLMs, the authors distinguish between evidence selection and interpretation, finding that most context spillover arises from differing interpretations under various roles rather than from retrieving different evidence. They evaluate two mitigation strategies—expressing investor mindset as a user profile instead of an assistant role, and separating evidence-based from personalized outputs—both of which reduce but do not eliminate spillover, with effectiveness varying across models.

By Ahmed Asaad, Amr Mohamed, Yang Zhang, Omneya Abdelsalam
arXiv AI
Aug 20

FinSkillBench: Evaluating AI Agents and Domain Skills for Investment Management

FinSkillBench is an evaluation suite that tests whether language model agents can use financial domain skills to solve investment management tasks across portfolio construction, risk management, and fundamental analysis. The benchmark contains 12 subtasks with 2,603 episodes, each providing point‑in‑time inputs, hidden ground truth, and a verifier. Experiments show that curated skill packages improve performance significantly, while self‑generated skills offer little benefit, indicating that reliable procedural skills are crucial for effective AI agents in this domain.

By Jermyn Zhen Yong Bek, Zhuang Qiang Bok, Zhongtian Sun
arXiv AI
Sep 21

Trustworthy FinAInce: Unpacking How AI-Mediated Financial Advice is Judged

The study investigates how people evaluate AI-generated financial advice by conducting a randomized vignette experiment with 285 U.S. adults. Participants were presented with consistent financial recommendations delivered in three styles—AI, expert, and online community—alongside source labels. The results show that advice style most strongly influenced message and safety appraisals, expert labels increased perceived source knowledge, and decision context shaped risk and safety judgments, with these appraisals explaining a large portion of overall quality, trust, and intended reliance.

By Aryan Ramchandra Kapadia, Eshwar Chandrasekharan, Koustuv Saha